EDBT 2026 Demo / reviewers in the wild / expert
Yuan Zeng 0003
dblp:120/6940-3
· DBLP profile ↗
3ranked-venue papers
1as first author
0since 2021 · last 2020
0000-0002-5550-9379ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Memory systems · 94% Energy-efficient computing · 6% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache |
0.4 | 1 | 2020 | RnR: A Software-Assisted Record-and-Replay Hardware Prefetcher · MICRO 2020 |
Memory systems
cache design |
0.4 | 1 | 2020 | Scrabble: A Fine-Grained Cache with Adaptive Merged Block · IEEE Trans. Computers 2020 |
Memory systems › cache › prefetching
hardware prefetching |
0.4 | 1 | 2020 | RnR: A Software-Assisted Record-and-Replay Hardware Prefetcher · MICRO 2020 |
Memory systems
memory access latency |
0.4 | 1 | 2020 | RnR: A Software-Assisted Record-and-Replay Hardware Prefetcher · MICRO 2020 |
Memory systems › cache
cache miss |
0.1 | 1 | 2020 | RnR: A Software-Assisted Record-and-Replay Hardware Prefetcher · MICRO 2020 |
Energy-efficient computing › low-power design
on-chip power reduction |
0.1 | 1 | 2020 | Scrabble: A Fine-Grained Cache with Adaptive Merged Block · IEEE Trans. Computers 2020 |
Methods — techniques the papers use, named apart from their topics
software-assisted hardware prefetching · 0.9record and replay · 0.9tag sharing · 0.4adaptive block merging · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | RnR: A Software-Assisted Record-and-Replay Hardware PrefetcherabstractApplications with irregular memory access patterns do not benefit well from the memory hierarchy as applications that have good locality do. Relatively high miss ratio and long memory access latency can cause the processor to stall and degrade system performance. Prefetching can help to hide the miss penalty by predicting which memory addresses will be accessed in the near future and issuing memory requests ahead of the time. However, software prefetchers add instruction overhead, whereas hardware prefetchers cannot efficiently predict irregular memory access sequences with high accuracy. Fortunately, in many important irregular applications (e.g., iterative solvers, graph algorithms, and sparse matrix-vector multiplication), memory access sequences repeat over multiple iterations or program phases. When the patterns are long, a conventional spatial-temporal prefetcher can not achieve high prefetching accuracy, but these repeating patterns can be identified by programmers.In this work, we propose a software-assisted hardware prefetcher that focuses on repeating irregular memory access patterns for data structures that cannot benefit from conventional hardware prefetchers. The key idea is to provide a programming interface to record cache miss sequence on the first appearance of a memory access pattern and prefetch through replaying the pattern on the following repeats. The proposed Record-and-Replay (RnR) prefetcher provides a lightweight software interface so that the programmers can specify in the application code: 1) which data structures have irregular memory accesses, 2) when to start the recording, and 3) when to start the replay (prefetching). This work evaluated three irregular workloads with different inputs. For the evaluated workloads and inputs, the proposed RnR prefetcher can achieve on average 2.16× speedup for graph applications and 2.91× speedup for an iterative solver with a sparse matrix-vector multiplication kernel. By leveraging the knowledge from the programmers, the proposed RnR prefetcher can achieve over 95% prefetching accuracy and miss coverage. Chao Zhang 0039, Yuan Zeng 0003, John Shalf |
MICRO | 2 |
| 2020 | Scrabble: A Fine-Grained Cache with Adaptive Merged BlockabstractA large fraction of the microprocessor energy is consumed by the data movement in the system. One of the reasons is the inefficiency in the conventional cache design. Cache blocks larger than a word are used in conventional caches to exploit spatial locality. However, many applications only use a small part of a cache block before its eviction. Transferring and storing unused data wastes bandwidth, energy, and limited cache space. Prior work on fine-grained caches can reduce data access and storage granularity to reduce the amount of unused data. However, small data blocks typically require greater metadata and control overhead. Sharing the common bits among tags of fine-grained blocks can reduce the metadata overhead but the constraints on which fine-grained blocks can share tag bits can cause fragmentation. This work proposes scrabble, a fine-grained cache that can merge multiple non-contiguous fine-grained blocks into a variable size merged block. The length of the shared tag is maximized to reduce the metadata overhead. The space utilization is improved by supporting merged blocks with variable size. The control overhead can be reduced by moving the merged block together from memory to the last level cache. For applications with poor spatial locality, Scrabble cache can achieve more than 40 percent of performance improvement. Even for application with good spatial locality, the speedup is still more than 7 percent. In general, for an evaluated set of benchmarks, Scrabble cache achieves an average of 2.41× effective capacity over the baseline cache with the same cache capacity which leads to a 16.7 percent performance improvement and an 11 percent on-chip energy reduction. As compared to a state-of-the-art fine-grained cache, Scrabble cache achieves a 1.25× effective capacity, a 7.9 percent speedup, and a 5.8 percent on-chip energy reduction. Chao Zhang 0039, Yuan Zeng 0003 |
IEEE Trans. Computers | 2 |
| 2018 | A Supervised Stdp-Based Training Algorithm for Living Neural NetworksabstractNeural networks have shown great potential in many applications like speech recognition, drug discovery, image classification, and object detection. Neural network models are inspired by biological neural networks, but they are optimized to perform machine learning tasks on digital computers. The proposed work explores the possibility of using living neural networks in vitro as the basic computational elements for machine learning applications. A new supervised STDP-based learning algorithm is proposed in this work, which considers neuron engineering constraints. A 74.7% accuracy is achieved on the MNIST benchmark for handwritten digit recognition. Yuan Zeng 0003, Kevin Devincentis, Zubayer Ibne Ferdous, Zhiyuan Yan 0001, Yevgeny Berdichevsky |
ICASSP | 1 |