VLDB 2026 Research / reviewers in the wild / expert
Yun Wang 0039
dblp:36/3235-39
· DBLP profile ↗
19ranked-venue papers
3as first author
17since 2021 · last 2026
0009-0009-7969-6891ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 1 first-author · 10 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RealSanitizer: Detecting Floating-Point Errors via a Dynamic Precision Oracle
Yuantao Hu, Shizhong Zhao, Yun Wang 0039, Mingyuan Xia 0001 |
TASE | 3 |
| 2026 | SpiderSense: Lightweight Last-Level Cache Management via Time Period Tagging for LLC-Critical WorkloadsabstractMulti-tenant clouds enhance resource sharing among Virtual Machines (VMs) to boost overall utilization and reduce power consumption. However, this also introduces interference among workloads from different tenants and impedes VM performance isolation. In this article, we first demonstrate that the last-level cache (LLC) in CPUs, which is inherently shared by all VMs on the same physical machine, becomes a significant contending resource for LLC-critical workloads, leading to notable performance imbalances under the default hardware caching strategy. Although recent studies on LLC scheduling have progressed, they often require detailed profiling of user workloads or rely on hyperparameter tuning, limiting their applicability to private clusters or specific scenarios. We propose SpiderSense, a software-initiated LLC partitioner for managing Virtual Machine Monitors (VMM), to address these limitations. SpiderSense leverages modern yet off-the-shelf server CPU features to adaptively orchestrate LLC allocation among running black-boxed user VMs. SpiderSense dynamically samples VMs and calculates their fair share of LLC to allocate them while fully improving performance isolation among VMs. We experiment with SpiderSense using typical LLC-critical workloads, representative of the types of applications that stress LLC performance, such as Memcached and Llama. Our results show that SpiderSense improves performance by up to 40% in numerous colocation scenarios compared to current solutions. Zhixiang Wei, Zhibai Huang, James Yen, Tianlei Xiong, Kailiang Xu, Yucheng Zheng, Xingzi Yu, Yun Wang 0039, Zhengwei Qi |
ACM Trans. Archit. Code Optim. | 9 |
| 2026 | gPooling: An Elastic GPU Resource Management Framework for On-Demand Virtualization in Shared Accelerator Clusters
Kaicheng Guo, Chen Chen 0067, Yun Wang 0039, Pengwei Du, Zhengwei Qi, Haibing Guan |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | Design and Operation of Elastic GPU-Pooling on Campus
Kaicheng Guo, Yun Wang 0039, Semakin Anton, Tovmachenko Dmitry, Jiajie Sheng, Jianwen Wei, James Lin 0001, Zhengwei Qi, Haibing Guan |
Euro-Par (1) | 3 |
| 2025 | DevTrace: Lightweight Plug-In Design for PCIe Transaction Tracing in Edge Intelligence WorkloadsabstractThe complexity of host-peripheral interactions during high-load tasks poses significant challenges for system optimization, with existing tracing tools degrading performance by up to 5.39×. We introduce DevTrace, a novel low-overhead tracing framework for peripheral interactions. Its modular architecture separates data collection from kernel-level operations, enabling lightweight tracing with minimal driver modifications across entire classes of devices. By eliminating heavy kernel tracing interrupts, DevTrace reduces overhead to negligible levels while maintaining data accuracy. In edge-based intelligence deployments, DevTrace achieves a 128× reduction in memory usage and approximately 10× lower CPU overhead compared to page-fault-based solutions. It significantly reduces data loss and performance degradation under high-load conditions, establishing it as a reliable tool for analyzing host-peripheral interactions and optimizing performance in resource-constrained environments. We also discuss potential extensions to eBPF to further decouple tracing from driver frameworks. Zhibai Huang, Kailiang Xu, Zhixiang Wei, Yinghao Deng, Chen Chen 0067, Yun Wang 0039, Fangxin Liu, Mingyuan Xia 0001, Zhengwei Qi |
ICCAD | 6 |
| 2025 | gFlow: Distributed Real-Time Reverse Remote Rendering System Model
Yixiao Xu, Wanzhao Xu, Yicheng Gu, Yun Wang 0039, Jiangyuan Ma, Zhengwei Qi |
MMM (2) | 5 |
| 2025 | ARMing x86 Games: Accelerating Binary Translation Using Software-Only Validated Flag Speculation
James Yen, Zhibai Huang, Zhixiang Wei, Chen Chen 0067, Senhao Yu, Yun Wang 0039, Hao Wang 0022, Zhengwei Qi |
MobiSys | 8 |
| 2025 | To PRI or Not To PRI, That's the question
Yun Wang 0039, Xianting Tian, Ben Luo, Zhixiang Wei, Zhibai Huang, Kailiang Xu, Kaihuan Peng, Kaijie Guo, Guangjian Wang, Shengdong Dai, Yibin Shen, Jiesheng Wu, Zhengwei Qi |
OSDI | 1 |
| 2025 | Effectively Virtual Page Prefetching via Spatial-Temporal Patterns for Memory-intensive Cloud ApplicationsabstractIn today's data-driven era, the explosive growth of global data volume has led to an increasing consumption of computing and storage resources. Effective management of virtual machines (VMs) memory usage is critical for cloud vendors to optimize system performance and resource utilization. Existing memory prefetching methods often slow down system performance, creating a difficult balance between maintaining service quality and optimizing resource use. For instance, Leap, which primarily utilizes address information, performs poorly in VM environments. The main issue is the performance drop caused by the reuse of memory resources in virtualized environments, a common situation in public clouds. Yun Wang 0039, Tianmai Deng, Ben Luo, Yibin Shen, Zhixiang Wei, Yixiao Xu, Minglang Huang, Zhengwei Qi |
PPoPP | 1 |
| 2025 | gCom: Fine-grained Compressors in Graphics Memory of Mobile GPUabstractToday, GPUs significantly boost rendering performance. However, the high memory requirements limit their use, especially on low-end mobile platforms. Compression techniques have been widely adopted to reduce memory consumption but face two primary issues when applied to mobile GPUs: (1) low repetition ratio caused by small raw data sizes and concurrency, and (2) low locality caused by unpredictable rendering behaviors. These two limitations result in a low compression ratio when compressors are applied to low-end mobile devices. This article introduces gCom , a fine-grained rendering compressor accelerated by GPUs. To improve the compression ratio, gCom incorporates the following innovations. First, unlike other compression techniques that use frames or tiles as basic processing units, gCom is the first to employ a fine-grained processing unit (i.e., the color channel), enhancing repetition amplification without increasing raw data. Second, gCom introduces two key features— Hierarchical Delta and Channel Decorrelator —which maximize the locality of adjacent channels and reduce raw data size. Third, to maintain the original GPU throughput, gCom revolutionizes the Golomb-Rice algorithm and proposes a new compression approach, the Parallel-Oriented Golomb-Rice algorithm, enabling parallel execution of both decompression and compression processes. The entire design of gCom utilizes only idle resources and existing commands on mobile GPUs, thus keeping purchasing costs low. To date, gCom has improved the channel locality by nearly 50%. The best compression achievement received by gCom has reached around 20%. Dongjie Tang, Yun Wang 0039, Yicheng Gu, Fangxin Liu, Zhengwei Qi |
ACM Trans. Archit. Code Optim. | 3 |
| 2025 | Rethinking Virtual Machines Live Migration for Memory DisaggregationabstractResource underutilization has troubled data centers for several decades. On the CPU front, live migration plays a crucial role in reallocating CPU resources. Nevertheless, contemporary Virtual Machine (VM) live migration methods are burdened by substantial resource consumption. In terms of memory management, disaggregated memory offers an effective solution to enhance memory utilization, but leaves a gap in addressing CPU underutilization. Our findings highlight a considerable opportunity to optimize live migration in the context of disaggregated memory systems. We introduce Anemoi, a resource management system that seamlessly integrates VM live migration with memory disaggregation to address the aforementioned gap. In the context of disaggregated memory, remote memory becomes accessible from destination nodes, effectively eliminating the need for extensive network transmission of memory pages, and thereby significantly reducing migration time. In addition, we propose using memory replicas as an optimization to the live migration system. To mitigate the overhead of potential excessive memory consumption, we develop a dedicated compression algorithm. Our evaluations demonstrate that Anemoi leads to a notable 69% reduction in network bandwidth utilization and an impressive 83% reduction in migration time compared to traditional VM live migration. Additionally, our compression algorithm achieves an outstanding space-saving rate of 83.6%. Xingzi Yu, Xingguo Jia, Yun Wang 0039, Senhao Yu, Zhengwei Qi |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | Sudoku: Scalable High-Density Cloud Rendering Multi-Client Architecture
Yun Wang 0039, Bing Deng, Xia Jiang, Xuyan Hu, Dongjie Tang, Randy Xu, Yijin Sun, Zhengwei Qi |
IEEE Trans. Serv. Comput. | 1 |
| 2024 | gVulkan: Scalable GPU Pooling for Pixel-Grained Rendering in Ray Tracing
Yicheng Gu, Yun Wang 0039, Yunfan Sun, Yuxin Xiang, Xuyan Hu, Zhengwei Qi, Haibing Guan |
USENIX ATC | 2 |
| 2023 | User-guided Page Merging for Memory Deduplication in Serverless SystemsabstractServerless computing is an emerging cloud paradigm that offers an elastic and scalable allocation of computing resources with pay-as-you-go billing. In the Function-as-a-Service (FaaS) programming model, applications comprise short-lived and stateless serverless functions executed in isolated containers or microVMs, which can quickly scale to thousands of instances and process terabytes of data. This flexibility comes at the cost of duplicated runtimes, libraries, and user data spread across many function instances, and cloud providers do not utilize this redundancy. The memory footprint of serverless forces removing idle containers to make space for new ones, which decreases performance through more cold starts and fewer data caching opportunities.We address this issue by proposing deduplicating memory pages of serverless workers with identical content, based on the content-based page-sharing concept of Linux Kernel Same-page Merging (KSM). We replace the background memory scanning process of KSM, as it is too slow to locate sharing candidates in short-lived functions. Instead, we design User-Guided Page Merging (UPM), a built-in Linux kernel module that leverages the madvise system call: we enable users to advise the kernel of memory areas that can be shared with others. We show that UPM reduces memory consumption by up to 55% on 16 concurrent containers executing a typical image recognition function, more than doubling the density for containers of the same function that can run on a system. Marcin Copik, Yun Wang 0039, Alexandru Calotoiu, Torsten Hoefler |
IEEE Big Data | 3 |
| 2023 | Rethinking Virtual Machines Live Migration for Memory DisaggregationabstractResource underutilization has troubled data centers for several decades. Memory disaggregation provides an efficient way to improve memory utilization while leaving a missing puzzle piece on CPU underutilization. Live migration is an essential method for CPU resource reallocation. However, the state-of-the-art Virtual Machines (VM) live migration suffers from significant resource consumption. We discover the substantial potential for optimizing live migration in disaggregated memory systems. We propose Anemoi, a source management system incorporating VM live migration into memory disaggregation to fill in the missing piece. Disaggregated memory enables the read-only replica to be accessible from destination nodes, eliminating considerable network transmission for memory pages and saving migration time. Regarding the potential overwhelming memory consumption of duplicate read-only replicas, we design a dedicated compression algorithm. The evaluation shows that Anemoi reduces the network bandwidth use and the migration time by 69% and 83%, respectively, compared to VM live migration. The compression can achieve a space-saving of 83.6%. Xingguo Jia, Xingzi Yu, Yun Wang 0039, Senhao Yu, Zhengwei Qi |
CLUSTER | 3 |
| 2023 | Bindox: An Efficient and Secure Cross-System IPC Mechanism for Multi-Platform ContainersabstractContainerization is widely used for isolation in various applications because it is lightweight, scalable, and portable.In modern distributed systems, seamless inter-process communication (IPC) between multi-platform containers is essential for a range of applications and services, including microservices, cloud computing, and Internet of Things (IoT) devices.However, secure and efficient communication between containers on the same host is challenging, especially when different operating systems are involved.This paper introduces Bindox, a lightweight, efficient, and secure IPC mechanism that enables seamless communication across multiple platforms, including Android and Linux.Bindox uses shared memory for data transfer and implements a stable client-server architecture, ensuring high performance and ease of maintenance.Additionally, Bindox provides a robust security mechanism that guarantees confidentiality, integrity, and availability of the communication channel.Experimental results demonstrate that Bindox outperforms existing networking and IPC methods in terms of memory use, latency, and CPU usage, making it a promising solution for efficient and secure communication between multi-platform containers. Yuxin Xiang, Bing Deng, Randy Xu, Marc Mao, Yun Wang 0039, Zhengwei Qi |
SEKE | 6 |
| 2023 | rShare: Alleviating long startup on the Cloud-rendering platform through de-systemization
Dongjie Tang, Marc Mao, Cathy Bao, Qiming Shi, Randy Xu, Mohammad R. Haghighat, Yun Wang 0039, Zhengwei Qi, Haibing Guan, Xiaojie Cao |
J. Syst. Archit. | 9 |
| 2020 | gRemote: API-Forwarding Powered Cloud RenderingabstractTraditional GPU resource allocation approaches, widely adopted in today's data centers, only focus on the server-side functions while ignoring the client-side. These approaches waste client-side hardware resources. To solve this problem, remote API-forwarding architectures appear. Through running applications on the client-side, remote API-forwarding architectures offload some workloads to the client. However, many remote API-forwarding systems suffer from one big issue: shared-resource interference, stemming from two reasons: (a) GPU resource racing caused by resource overuse for a single client, and (b) CPU resource racing caused by resource shortage among clients. This paper presents gRemote, an open-source GPU-remoting system that can address this issue. To mitigate the CPU resource shortage, gRemote improves CPU configurations by expanding CPU resources from the server-side to both server- and client-side. To maintain the reasonable GPU usage for individual tasks, we innovate a new resource-sharing mechanism called GPU throttle. gRemote supports 1,228 OpenGL commands with around 10% shared-resource interference. Dongjie Tang, Yun Wang 0039, Linsheng Li, Jiacheng Ma 0001, Xue (Steve) Liu, Zhengwei Qi, Haibing Guan |
HPDC | 2 |
| 2019 | A distributed hypervisor for resource aggregation: posterabstractScale-out has become the standard answer to data analysis, machine learning and many other fields. Contrary to common belief, scale-up machines can outperform scale-out clusters for a considerable portion of tasks. However, those scale-up machines are not economical and may not be affordable for small businesses. This paper presents GiantVM, a distributed hypervisor that aggregates resources from multiple physical machines, providing the guest OS with a uniform hardware abstraction. We propose techniques to deal with the challenges of CPU, Memory, and I/O virtualization in distributed environments. Yubin Chen, Zhuocheng Ding, Yun Wang 0039, Zhengwei Qi, Haibing Guan |
PPoPP | 4 |