Yunmo Zhang

dblp:334/2127 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-7462-2780ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 ANG: Accelerating NFA processing on GPUs via Exploring Multi-Level Fine-Grained Parallelism
abstract
Finite Automata (FA) processing is a core computation in various real-world applications. Over the past decades, extensive efforts have been dedicated to accelerating FA processing on modern parallel platforms, particularly GPUs, due to their high memory bandwidth and massive hardware parallelism. As Non-deterministic Finite Automata (NFA)-based applications have strong and growing demands for real-time data analytics nowadays, reducing latency in automata processing has become a critical priority. However, existing approaches face significant challenges when limited parallelism is exposed in NFA computations. In this work, we explore opportunities of introducing fine-grained parallelism from various sources and addressing the limitations of fast NFA processing. Specifically, by analyzing different NFA parallelization schemes, we identify the major performance issue caused by insufficient state-level parallelism in conventional designs. To overcome the bottleneck, this work introduces speculative parallelization tailored for GPU-based NFA processing, thus effectively exploiting fine-grained parallelism across multilevels, with a particular focus on input-chunk-level parallelism. To realize speculative parallelization in practice, we develop $A N G$, a latency-oriented NFA processing framework that overcomes key implementation challenges on GPUs. We evaluate the efficiency of ANG on a set of representative NFAs with diverse properties. Experimental results demonstrate that ANG achieves significant performance improvement compared to state-of-theart techniques, with reaching $11.74 \times$ speedup on average (and up to $49.88 \times$ in extreme cases).
Yuguang Wang 0005, Yunmo Zhang, Junqiao Qiu, Zhenlin Wang 0003
PACT2
2025 PIE: Enabling Fast and Scalable Incremental Evolving Graph Analytics on Persistent Memory
abstract
Graph processing is crucial for unstructured-data-driven applications in various domains.In recent years, there has been a growing need to perform real-time analytics on largescale evolving graphs, which involves evaluating a graph query on a sequence of snapshots within a given time window.Some prior studies have explored utilizing persistent memory (PM) technologies, such as non-volatile memory, for efficient evolving graph analytics.However, the latest incremental processing designs fail to fully exploit the PM potential, suffering from severe read and write amplification during update ingestion and query evaluation.In this paper, we develop PIE, a PM-based incremental processing framework for fast and scalable evolving graph analytics.We first observe that leveraging CommonGraph, a recently proposed DRAM-based incremental approach that transforms costly deletions into additions, can significantly improve efficiency for evolving graph analytics in PM, although the direct adaptation introduces significant PM access inefficiencies.To enable PM-friendly incremental processing, PIE introduces a logical graph view abstraction that is detached from the physical storage to avoid extra PM writes, and a
Yunmo Zhang, Jiacheng Huang 0002, Xizhe Yin, Junqiao Qiu, Hong Xu 0001, Chun Jason Xue
ICS1
2025 Inferring Likely Counting-related Atomicity Program Properties for Persistent Memory
Yunmo Zhang, Junqiao Qiu, Hong Xu 0001, Chun Jason Xue
USENIX ATC1
2024 More Apps, Faster Hot-Launch on Mobile Devices via Fore/Background-aware GC-Swap Co-design
abstract
Faster app launching is crucial for the user experience on mobile devices. Apps launched from a background cached state, called hot-launching, have much better performance than apps launched from scratch. To increase the number of hot-launches, leading mobile vendors now cache more apps in the background by enabling swap. Recent work also proposed reducing the Java heap to increase the number of cached apps. However, this paper found that existing methods deteriorate app hot-launch performance while increasing the number of cached apps. To simultaneously improve the number of cached apps and hot-launch performance, this paper proposes Fleet, a foreground/background-aware GC-swap co-design framework. To enhance app-caching capacity, Fleet limits the tracing range of GC to background objects only, avoiding touching long-lifetime foreground objects. To improve hot-launch performance, Fleet identifies objects that will be accessed during the next hot-launch and uses runtime information to guide the swap scheme in the OS. In addition, Fleet aggregates small objects with similar access patterns into the same pages to improve swap efficiency. We implemented Fleet in AOSP and evaluated its performance with different types of apps. Experimental results show that Fleet achieves a 1.59× faster hot-launch time and caches 1.21× more apps than Android.
Jiacheng Huang 0002, Yunmo Zhang, Junqiao Qiu, Yu Liang 0004, Rachata Ausavarungnirun, Qing'an Li, Chun Jason Xue
ASPLOS (3)2
2022 Probabilistic Analysis of Network Availability
abstract
In recent years, quantitative verification concerning network availability has received increasing attention due to its practical importance in network management. Existing work extends upon qualitative verification and tries to answer whether a network would have any link overload under failures and dynamic traffic. The simple yes-or-no question falls short of accurately characterizing the robustness of a network under uncertainties, which the operators are in dire need of. Thus, we argue that it is necessary to design a probabilistic framework to analyze network availability comprehensively. We propose Pita, a novel network analysis framework that outputs the overall probability of the network being unavailable under a range of failure scenarios and traffic demands. We formalize the problem and show that it is #P-hard which does not admit deterministic approximation solutions. We further develop an improved randomized approximation that exploits the structural property of our problem to reduce the computational cost of the so-called boundary oracle procedure, a key bottleneck of the approximation, without any accuracy loss. Evaluation with real topologies shows that Pita provides up to 2.25x speedup over state-of-the-art solutions, and can be effectively used in many network management tasks such as identifying high-risk failure scenarios, and aiding robust traffic engineering design.
Yunmo Zhang, Hong Xu 0001, Chun Jason Xue, Tei-Wei Kuo
ICNP1