Yanpeng Yu

dblp:296/3723 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0002-8292-935XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 CORD: Low-Latency, Bandwidth-Efficient and Scalable Release Consistency via Directory Ordering
abstract
Increasingly, multi-processing unit (PU) systems (e.g., CPU-GPU, multi-CPU, multi-GPU, etc.) are embracing cache-coherent shared memory to facilitate inter-PU communication.The coherence protocols in these systems support write-through accesses that place the data directly at the LLC to enable efficient producer-consumer communications pervasive in AI/ML workloads.Moreover, release consistency has emerged as the standard memory model in such systems due to its programming simplicity and ability to support high performance.In today's multi-PU systems, the source processor that issues the writes also orders them to enforce release consistency, even for write-through accesses.Unfortunately, such source ordering of write-through operations results in unnecessary communications between the source processor and the LLC directory, incurring significant performance, interconnect traffic, and energy overheads for multi-PU applications.To eliminate such communication, we present cord 1 , a novel cache coherence protocol that orders write-through accesses directly at the cache directory.cord employs several novel mechanisms to minimize the metadata required for ordering traffic while efficiently scaling to multiple directories.Evaluations atop the gem5 simulator show that compared to source ordering, cord improves application performance by 24% and reduces traffic by 13% on average while incurring < 1% storage, area, and power overheads.Compared to hand-optimized message-passing implementations, cord observes a mere 3% performance overhead and 6% more traffic on average with a significantly simpler programming model.
Yanpeng Yu, Nicolai Oswald, Anurag Khandelwal
ISCA1
2023 An Efficient Data Structure for Dynamic Graph on GPUs
abstract
There is a growing interest to offload dynamic graph computation to GPU and resort to its high parallel processing ability and larger memory bandwidths compared with CPUs. The existing GPU graph systems usually use compressed sparse row (CSR) as the de-facto structure. However, CSR has a critical weakness for dynamic change due to the large overhead of re-balance process after update. GPMA+ is a state-of-art dynamic PMA-based structure that uses PMA structure and segment-oriented parallel update procedure to address the dynamic weakness of CSR, but it still has a bottleneck on the array expansion. In this paper, we propose an leveled structure (called LPMA) instead of continue array to retain low time complexity and high parallel update and lift the expansion bottleneck of GPMA+. More specifically, we propose a series of optimization techniques, including bottom-up update, top-down update and on-demand hybrid update strategies as well as consistence-guaranteed parallel processing for update-query mixed workloads. We theoretically analyze the benefits of LPMA compared in terms of re-balance cost during updates. Extensive experiments on four large real-life graphs prove the superiority of LPMA compared with the-state-of-arts.
Lei Zou 0001, Fan Zhang 0050, Yinnian Lin, Yanpeng Yu
IEEE Trans. Knowl. Data Eng.4
2021 MIND: In-Network Memory Management for Disaggregated Data Centers
abstract
Memory disaggregation promises transparent elasticity, high resource utilization and hardware heterogeneity in data centers by physically separating memory and compute into network-attached resource "blades". However, existing designs achieve performance at the cost of resource elasticity, restricting memory sharing to a single compute blade to avoid costly memory coherence traffic over the network.
SeungSeob Lee, Yanpeng Yu, Yupeng Tang, Anurag Khandelwal, Lin Zhong 0001, Abhishek Bhattacharjee
SOSP2
2021 LPMA - An Efficient Data Structure for Dynamic Graph on GPUs
Fan Zhang 0050, Lei Zou 0001, Yanpeng Yu
WISE (1)3