Yuan Zeng 0003

dblp:120/6940-3 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
0since 2021 · last 2020
0000-0002-5550-9379ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 94% Energy-efficient computing · 6%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache
0.412020
RnR: A Software-Assisted Record-and-Replay Hardware Prefetcher · MICRO 2020
Memory systems
cache design
0.412020
Scrabble: A Fine-Grained Cache with Adaptive Merged Block · IEEE Trans. Computers 2020
Memory systems › cache › prefetching
hardware prefetching
0.412020
RnR: A Software-Assisted Record-and-Replay Hardware Prefetcher · MICRO 2020
Memory systems
memory access latency
0.412020
RnR: A Software-Assisted Record-and-Replay Hardware Prefetcher · MICRO 2020
Memory systems › cache
cache miss
0.112020
RnR: A Software-Assisted Record-and-Replay Hardware Prefetcher · MICRO 2020
Energy-efficient computing › low-power design
on-chip power reduction
0.112020
Scrabble: A Fine-Grained Cache with Adaptive Merged Block · IEEE Trans. Computers 2020

Methods — techniques the papers use, named apart from their topics

software-assisted hardware prefetching · 0.9record and replay · 0.9tag sharing · 0.4adaptive block merging · 0.4
YearPublicationVenuePosition
2020 RnR: A Software-Assisted Record-and-Replay Hardware Prefetcher
abstract
Applications with irregular memory access patterns do not benefit well from the memory hierarchy as applications that have good locality do. Relatively high miss ratio and long memory access latency can cause the processor to stall and degrade system performance. Prefetching can help to hide the miss penalty by predicting which memory addresses will be accessed in the near future and issuing memory requests ahead of the time. However, software prefetchers add instruction overhead, whereas hardware prefetchers cannot efficiently predict irregular memory access sequences with high accuracy. Fortunately, in many important irregular applications (e.g., iterative solvers, graph algorithms, and sparse matrix-vector multiplication), memory access sequences repeat over multiple iterations or program phases. When the patterns are long, a conventional spatial-temporal prefetcher can not achieve high prefetching accuracy, but these repeating patterns can be identified by programmers.In this work, we propose a software-assisted hardware prefetcher that focuses on repeating irregular memory access patterns for data structures that cannot benefit from conventional hardware prefetchers. The key idea is to provide a programming interface to record cache miss sequence on the first appearance of a memory access pattern and prefetch through replaying the pattern on the following repeats. The proposed Record-and-Replay (RnR) prefetcher provides a lightweight software interface so that the programmers can specify in the application code: 1) which data structures have irregular memory accesses, 2) when to start the recording, and 3) when to start the replay (prefetching). This work evaluated three irregular workloads with different inputs. For the evaluated workloads and inputs, the proposed RnR prefetcher can achieve on average 2.16× speedup for graph applications and 2.91× speedup for an iterative solver with a sparse matrix-vector multiplication kernel. By leveraging the knowledge from the programmers, the proposed RnR prefetcher can achieve over 95% prefetching accuracy and miss coverage.
Chao Zhang 0039, Yuan Zeng 0003, John Shalf
MICRO2
2020 Scrabble: A Fine-Grained Cache with Adaptive Merged Block
abstract
A large fraction of the microprocessor energy is consumed by the data movement in the system. One of the reasons is the inefficiency in the conventional cache design. Cache blocks larger than a word are used in conventional caches to exploit spatial locality. However, many applications only use a small part of a cache block before its eviction. Transferring and storing unused data wastes bandwidth, energy, and limited cache space. Prior work on fine-grained caches can reduce data access and storage granularity to reduce the amount of unused data. However, small data blocks typically require greater metadata and control overhead. Sharing the common bits among tags of fine-grained blocks can reduce the metadata overhead but the constraints on which fine-grained blocks can share tag bits can cause fragmentation. This work proposes scrabble, a fine-grained cache that can merge multiple non-contiguous fine-grained blocks into a variable size merged block. The length of the shared tag is maximized to reduce the metadata overhead. The space utilization is improved by supporting merged blocks with variable size. The control overhead can be reduced by moving the merged block together from memory to the last level cache. For applications with poor spatial locality, Scrabble cache can achieve more than 40 percent of performance improvement. Even for application with good spatial locality, the speedup is still more than 7 percent. In general, for an evaluated set of benchmarks, Scrabble cache achieves an average of 2.41× effective capacity over the baseline cache with the same cache capacity which leads to a 16.7 percent performance improvement and an 11 percent on-chip energy reduction. As compared to a state-of-the-art fine-grained cache, Scrabble cache achieves a 1.25× effective capacity, a 7.9 percent speedup, and a 5.8 percent on-chip energy reduction.
Chao Zhang 0039, Yuan Zeng 0003
IEEE Trans. Computers2
2018 A Supervised Stdp-Based Training Algorithm for Living Neural Networks
abstract
Neural networks have shown great potential in many applications like speech recognition, drug discovery, image classification, and object detection. Neural network models are inspired by biological neural networks, but they are optimized to perform machine learning tasks on digital computers. The proposed work explores the possibility of using living neural networks in vitro as the basic computational elements for machine learning applications. A new supervised STDP-based learning algorithm is proposed in this work, which considers neuron engineering constraints. A 74.7% accuracy is achieved on the MNIST benchmark for handwritten digit recognition.
Yuan Zeng 0003, Kevin Devincentis, Zubayer Ibne Ferdous, Zhiyuan Yan 0001, Yevgeny Berdichevsky
ICASSP1