Chao Zhang 0039

dblp:94/3019-39 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
2since 2021 · last 2022
0000-0002-7892-5113ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 5 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Memory systems · 66% Hardware accelerators and domain-specific architectures · 13% Processor architecture and microarchitecture · 13%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture › instruction set architecture
ISA extension
0.612022
ASA: Accelerating Sparse Accumulation in Column-wise SpGEMM · ACM Trans. Archit. Code Optim. 2022
Hardware accelerators and domain-specific architectures › sparse matrix multiplication accelerator
sparse linear algebra accelerator
0.612022
ASA: Accelerating Sparse Accumulation in Column-wise SpGEMM · ACM Trans. Archit. Code Optim. 2022
Memory systems › memory hierarchy
cache hierarchy
0.512021
SPX64: A Scratchpad Memory for General-purpose Microprocessors · ACM Trans. Archit. Code Optim. 2021
Memory systems › on-chip memory
scratchpad memory
0.512021
SPX64: A Scratchpad Memory for General-purpose Microprocessors · ACM Trans. Archit. Code Optim. 2021
Memory systems
cache
0.412020
RnR: A Software-Assisted Record-and-Replay Hardware Prefetcher · MICRO 2020
Memory systems
cache design
0.412020
Scrabble: A Fine-Grained Cache with Adaptive Merged Block · IEEE Trans. Computers 2020
Memory systems › cache › prefetching
hardware prefetching
0.412020
RnR: A Software-Assisted Record-and-Replay Hardware Prefetcher · MICRO 2020
Memory systems
memory access latency
0.412020
RnR: A Software-Assisted Record-and-Replay Hardware Prefetcher · MICRO 2020
Parallel and multicore computing › parallel algorithms › parallel matrix algorithms
sparse general matrix-matrix multiplication
0.212022
ASA: Accelerating Sparse Accumulation in Column-wise SpGEMM · ACM Trans. Archit. Code Optim. 2022
Memory systems › cache
cache miss
0.112020
RnR: A Software-Assisted Record-and-Replay Hardware Prefetcher · MICRO 2020
Energy-efficient computing › low-power design
on-chip power reduction
0.112020
Scrabble: A Fine-Grained Cache with Adaptive Merged Block · IEEE Trans. Computers 2020

Methods — techniques the papers use, named apart from their topics

software-assisted hardware prefetching · 0.9record and replay · 0.9set-associative cache search · 0.6pipelined accumulation · 0.6virtual addressing · 0.5set-associative design · 0.5tag sharing · 0.4adaptive block merging · 0.4
YearPublicationVenuePosition
2022 ASA: Accelerating Sparse Accumulation in Column-wise SpGEMM
abstract
Sparse linear algebra is an important kernel in many different applications. Among various sparse general matrix-matrix multiplication (SpGEMM) algorithms, Gustavson’s column-wise SpGEMM has good locality when reading input matrix and can be easily parallelized by distributing the computation of different columns of an output matrix to different processors. However, the sparse accumulation (SPA) step in column-wise SpGEMM, which merges partial sums from each of the multiplications by the row indices, is still a performance bottleneck. The state-of-the-art software implementation uses a hash table for partial sum search in the SPA, which makes SPA the largest contributor to the execution time of SpGEMM. There are three reasons that cause the SPA to become the bottleneck: (1) hash probing requires data-dependent branches that are difficult for a branch predictor to predict correctly; (2) the accumulation of partial sum is dependent on the results of the hash probing, which makes it difficult to hide the hash probing latency; and (3) hash collision requires time-consuming linear search and optimizations to reduce these collisions require an accurate estimation of the number of non-zeros in each column of the output matrix. This work proposes ASA architecture to accelerate the SPA. ASA overcomes the challenges of SPA by (1) executing the partial sum search and accumulate with a single instruction through ISA extension to eliminate data-dependent branches in hash probing, (2) using a dedicated on-chip cache to perform the search and accumulation in a pipelined fashion, (3) relying on the parallel search capability of a set-associative cache to reduce search latency, and (4) delaying the merging of overflowed entries. As a result, ASA achieves an average of 2.25× and 5.05× speedup as compared to the state-of-the-art software implementation of a Markov clustering application and its SpGEMM kernel, respectively. As compared to a state-of-the-art hashing accelerator design, ASA achieves an average of 1.95× speedup in the SpGEMM kernel.
Chao Zhang 0039, Maximilian H. Bremer, Cy P. Chan, John Shalf
ACM Trans. Archit. Code Optim.1
2021 SPX64: A Scratchpad Memory for General-purpose Microprocessors
abstract
General-purpose computing systems employ memory hierarchies to provide the appearance of a single large, fast, coherent memory. In special-purpose CPUs, programmers manually manage distinct, non-coherent scratchpad memories. In this article, we combine these mechanisms by adding a virtually addressed, set-associative scratchpad to a general purpose CPU. Our scratchpad exists alongside a traditional cache and is able to avoid many of the programming challenges associated with traditional scratchpads without sacrificing generality (e.g., virtualization). Furthermore, our design delivers increased security and improves performance, especially for workloads with high locality or that interact with nonvolatile memory.
Shail Dave, Pantea Zardoshti, Robert Brotzman, Chao Zhang 0039, Aviral Shrivastava, Gang Tan, Michael F. Spear
ACM Trans. Archit. Code Optim.5
2020 ECC Cache: A Lightweight Error Detection for Phase-Change Memory Stuck-At Faults
abstract
DRAM scaling has been slowed down. Emerging non-volatile memories (e.g., Phase-Change Memory) promises higher density, better scalability, and persistence. However, endurance is a fundamental issue that hinders the broad adoption of PCM---after repeated writes, a PCM cell can get stuck at a value and be no longer programmable. The prevalence of this stuck-at fault issue requires error detection and correction mechanisms for PCM. Existing solutions such as verify-after-write adds additional latency to PCM writes, which degrades overall system performance. Other solutions like in-memory error-correcting code (ECC) requires a high storage overhead and introduces more reliability issue because ECC bits tend to wear out faster than the protected data bits.
Chao Zhang 0039, Khaled Abdelaal, Angel Chen, Xinhui Zhao, Wujie Wen
ICCAD1
2020 RnR: A Software-Assisted Record-and-Replay Hardware Prefetcher
abstract
Applications with irregular memory access patterns do not benefit well from the memory hierarchy as applications that have good locality do. Relatively high miss ratio and long memory access latency can cause the processor to stall and degrade system performance. Prefetching can help to hide the miss penalty by predicting which memory addresses will be accessed in the near future and issuing memory requests ahead of the time. However, software prefetchers add instruction overhead, whereas hardware prefetchers cannot efficiently predict irregular memory access sequences with high accuracy. Fortunately, in many important irregular applications (e.g., iterative solvers, graph algorithms, and sparse matrix-vector multiplication), memory access sequences repeat over multiple iterations or program phases. When the patterns are long, a conventional spatial-temporal prefetcher can not achieve high prefetching accuracy, but these repeating patterns can be identified by programmers.In this work, we propose a software-assisted hardware prefetcher that focuses on repeating irregular memory access patterns for data structures that cannot benefit from conventional hardware prefetchers. The key idea is to provide a programming interface to record cache miss sequence on the first appearance of a memory access pattern and prefetch through replaying the pattern on the following repeats. The proposed Record-and-Replay (RnR) prefetcher provides a lightweight software interface so that the programmers can specify in the application code: 1) which data structures have irregular memory accesses, 2) when to start the recording, and 3) when to start the replay (prefetching). This work evaluated three irregular workloads with different inputs. For the evaluated workloads and inputs, the proposed RnR prefetcher can achieve on average 2.16× speedup for graph applications and 2.91× speedup for an iterative solver with a sparse matrix-vector multiplication kernel. By leveraging the knowledge from the programmers, the proposed RnR prefetcher can achieve over 95% prefetching accuracy and miss coverage.
Chao Zhang 0039, Yuan Zeng 0003, John Shalf
MICRO1
2020 Scrabble: A Fine-Grained Cache with Adaptive Merged Block
abstract
A large fraction of the microprocessor energy is consumed by the data movement in the system. One of the reasons is the inefficiency in the conventional cache design. Cache blocks larger than a word are used in conventional caches to exploit spatial locality. However, many applications only use a small part of a cache block before its eviction. Transferring and storing unused data wastes bandwidth, energy, and limited cache space. Prior work on fine-grained caches can reduce data access and storage granularity to reduce the amount of unused data. However, small data blocks typically require greater metadata and control overhead. Sharing the common bits among tags of fine-grained blocks can reduce the metadata overhead but the constraints on which fine-grained blocks can share tag bits can cause fragmentation. This work proposes scrabble, a fine-grained cache that can merge multiple non-contiguous fine-grained blocks into a variable size merged block. The length of the shared tag is maximized to reduce the metadata overhead. The space utilization is improved by supporting merged blocks with variable size. The control overhead can be reduced by moving the merged block together from memory to the last level cache. For applications with poor spatial locality, Scrabble cache can achieve more than 40 percent of performance improvement. Even for application with good spatial locality, the speedup is still more than 7 percent. In general, for an evaluated set of benchmarks, Scrabble cache achieves an average of 2.41× effective capacity over the baseline cache with the same cache capacity which leads to a 16.7 percent performance improvement and an 11 percent on-chip energy reduction. As compared to a state-of-the-art fine-grained cache, Scrabble cache achieves a 1.25× effective capacity, a 7.9 percent speedup, and a 5.8 percent on-chip energy reduction.
Chao Zhang 0039, Yuan Zeng 0003
IEEE Trans. Computers1
2017 Enabling efficient fine-grained DRAM activations with interleaved I/O
abstract
DRAM contributes a significant part of the total system energy consumption, and row activation is one of the most energy inefficient components. Prior works on fine-grained DRAM activation rely on increasing the number of local wires to avoid degrading performance, which adds area overheads. This work proposes interleaved I/O to allow data transferring from different partially activated banks to share the global I/O. The proposed DRAM architecture allows half-, quarter-, or one-eighth- page activations without changing the wires. The system performance is competitive as compared with other fine-grained activation designs. For the evaluated benchmarks, an average of up to 15.7% performance improvement is achieved among all of the configurations. Furthermore, the total DRAM energy can be reduced by an average of 11.2% for halfpage, 17.2% for quarterpage, and 22.3% for one-eighth-page.
Chao Zhang 0039
ISLPED1