Yueyang Pan

dblp:313/3086 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Tolerate It if You Cannot Reduce It: Handling Latency in Tiered Memory
abstract
Current memory tiering systems mitigate asymmetric latency through page migration between tiers. While this approach effectively hides latency, it overlooks scenarios where latency could potentially be tolerated. We propose that an efficient system should integrate both latency reduction (migration) and latency tolerance strategies. Our research demonstrates the effectiveness of prefetchers in tolerating latency within such systems, highlighting their importance in the design of high-performance memory tiering solutions.
Musa Unal, Vishal Gupta 0006, Yueyang Pan, Yujie Ren, Sanidhya Kashyap
HotOS3
2025 Scalable Far Memory: Balancing Faults and Evictions
abstract
Page-based far memory systems transparently expand an application's memory capacity beyond a single machine without modifying application code. However, existing systems are tailored to scenarios with low application thread counts, and fail to scale on today's multi-core machines. This makes them unsuitable for data-intensive applications that both rely on far memory support and scale with increasing thread count. Our analysis reveals that this poor scalability stems from inefficient holistic coordination between page fault-in and eviction operations. As thread count increases, current systems encounter scalability bottlenecks in TLB shootdowns, page accounting, and memory allocation.
Yueyang Pan, Yash Lala, Musa Unal, Yujie Ren, SeungSeob Lee, Abhishek Bhattacharjee, Anurag Khandelwal, Sanidhya Kashyap
SOSP1
2025 Efficient Exact Resistance Distance Computation on Small-Treewidth Graphs: A Labelling Approach
abstract
Resistance distance computation is a fundamental problem in graph analysis, yet existing random walk-based methods are limited to approximate solutions and suffer from poor efficiency on small-treewidth graphs (e.g., road networks). In contrast, shortest-path distance computation achieves remarkable efficiency on such graphs by leveraging cut properties and tree decompositions. Motivated by this disparity, we first analyze the cut property of resistance distance. While a direct generalization proves impractical due to costly matrix operations, we overcome this limitation by integrating tree decompositions, revealing that the resistance distance r(s,t) depends only on labels along the paths from s and t to the root of the decomposition. This insight enables compact labelling structures. Based on this, we propose TreeIndex, a novel index method that constructs a resistance distance labelling of size O(n • h G ) in O(n • h G 2 • d max ) time, where h G (tree height) and d max (maximum degree) behave as small constants in many real-world small-treewidth graphs (e.g., road networks). Our labelling supports exact single-pair queries in O(h G ) time and single-source queries in O(n • h G ) time. Extensive experiments show that TreeIndex substantially outperforms state-of-the-art approaches. For instance, on the full USA road network, it constructs a 405 GB labelling in 7 hours (single-threaded) and answers exact single-pair queries in 10 -3 seconds and single-source queries in 190 seconds-the first exact method scalable to such large graphs.
Meihao Liao, Yueyang Pan, Rong-Hua Li 0001, Guoren Wang
Proc. ACM Manag. Data2
2024 Transparent Multicore Scaling of Single-Threaded Network Functions
abstract
This paper presents NFOS, a programming model, runtime, and profiler for productively developing software network functions (NFs) that scale on multicore machines. Writing shared-state concurrent systems that are both correct and scalable is still a serious challenge, which is why NFOS insulates developers from writing concurrent code.
Lei Yan 0003, Yueyang Pan, Diyu Zhou, George Candea, Sanidhya Kashyap
EuroSys2
2024 Monarch: A Fuzzing Framework for Distributed File Systems
Tao Lyu 0004, Zhiyao Feng, Yueyang Pan, Yujie Ren, Meng Xu 0001, Mathias Payer, Sanidhya Kashyap
USENIX ATC4
2023 Ship your Critical Section, Not Your Data: Enabling Transparent Delegation with TCLOCKS
Vishal Gupta 0006, Kumar Kartikeya Dwivedi, Yugesh Kothari, Yueyang Pan, Diyu Zhou, Sanidhya Kashyap
OSDI4
2023 Critique of "A Parallel Framework for Constraint-Based Bayesian Network Learning via Markov Blanket Discovery" by SCC Team From Peking University
abstract
Ankit Srivastava et al. (Srivastava et al. 2020) proposed a parallel framework for Constraint-Based Bayesian Network (BN) Learning via Markov Blanket Discovery (referred to as ramBLe) and implemented it over three existing BN learning algorithms, namely, GS, IAMB and Inter-IAMB. As part of the Student Cluster Competition at SC21, we reproduce the computational efficiency of ramBLe on our assigned Oracle cluster. The cluster has 4x36 cores in total with 100 Gbps RoCE v2 support and is equipped with CentOS-compatible Oracle Linux. Our experiments, covering the same three algorithms of the original ramBLe article (Srivastava et al. 2020), evaluate the strong and weak scalability of the algorithms using real COVID-19 data sets. We verify part of the conclusions from the original article and propose our explanation of the differences obtained in our results.Author: Please confirm or add details for any funding or financial support for the research of this article. ?>
Jiaqi Si, Junyi Guo, Zhewen Hao, Wenyang He, Yueyang Pan, Zhenxin Fu, Chun Fan 0001
IEEE Trans. Parallel Distributed Syst.6
2022 Critique of "MemXCT: Memory-Centric X-Ray CT Reconstruction With Massive Parallelization" by SCC Team From Peking University
abstract
Hidayetoluet al.(2019) proposed a novel memory-centric computation system, MemXCT. As a challenge at SC20, we reproduce the computational efficiency of MemXCT on our Azure cloud cluster. Our experiments evaluate the overall performance and the strong scalability with real datasets and verify part of the conclusions in the original article.
Zejia Fan, Zhewen Hao, Yueyang Pan, Pengcheng Xu 0005, Yuxuan Yan, Fangyuan Yang, Zhenxin Fu, Yun Liang 0001
IEEE Trans. Parallel Distributed Syst.4