EDBT 2026 Demo / reviewers in the wild / expert
Weiwei Jia 0001
dblp:61/9959-1
· DBLP profile ↗
16ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0001-5770-5472ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 6 first-author · 9 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Computer networks · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Optimizing Task Scheduling in Cloud VMs with Accurate vCPU AbstractionabstractThe paper shows that task scheduling in Cloud VMs hasn't evolved quickly to handle the dynamic vCPU resources. The existing vCPU abstraction cannot accurately depict the vCPU dynamics in capacity, activity, and topology, and these mismatches can mislead the scheduler, causing performance degradation and system anomalies. The paper proposes a novel solution, vSched, which probes accurate vCPU abstraction through a set of lightweight microbenchmarks (vProbers) without modifying the hypervisor, and leverages the probed information to optimize task scheduling in cloud VMs with three new techniques: biased vCPU selection, intra-VM harvesting, and relaxed work conservation. Our evaluation of vSched's implementation in x86 Linux Kernel demonstrates that it can effectively improve both system throughput and workload latency across various VM types in the dynamic multi-cloud environment. Edward Guo, Weiwei Jia 0001, Xiaoning Ding, Jianchen Shan |
EuroSys | 2 |
| 2025 | Rethinking Tiered Storage: Talk to File Systems, Not Device DriversabstractDifferent storage technologies motivate the development of specialized file systems tailored to specific device types. A tiered file system aggregates such device types into a single file system. We argue that the current practice of developing tiered file systems tends to lag behind that of device-specific file systems because, inherently, developers are burdened with addressing multiple device types simultaneously, rather than specializing. We propose to solve this problem using Mux, a new tiered file system that accesses different device types indirectly through device-specific file systems, rather than directly through device drivers. Despite introducing an additional indirection layer, we show that Mux significantly outperforms Strata, a research tiered file system, because it utilizes specialized production-ready file systems. Compared with direct access to per-device file systems (with no tiering), Mux adds a worst-case read latency overhead of 6.6% to 87.3%, and a write throughout overhead of 1.6% to 3.5% across devices. We contend that Mux's separation of tiering and specialization concerns enables progressive evolution and flexible integration of heterogeneous storage devices. Jiyuan Zhang 0003, Jongyul Kim 0001, Chloe Alverti, Peizhe Liu, Weiwei Jia 0001, Tianyin Xu |
HotOS | 5 |
| 2025 | EMT: An OS Framework for New Memory Translation Architectures
Siyuan Chai 0001, Jiyuan Zhang 0003, Jongyul Kim 0001, Fan Chung Graham, Jovan Stojkovic, Weiwei Jia 0001, Dimitrios Skarlatos 0002, Josep Torrellas, Tianyin Xu |
OSDI | 7 |
| 2024 | Direct Memory Translation for Virtualized CloudsabstractVirtual memory translation has become a key performance bottleneck of memory-intensive workloads in virtualized cloud environments. On the x86 architecture, a nested translation needs to sequentially fetch up to 24 page table entries (PTEs). This paper presents Direct Memory Translation (DMT), a hardware-software extension for x86-based virtual memory that minimizes translation overhead while maintaining backward compatibility with x86. In DMT, the OS manages last-level PTEs in a contiguous physical memory region, termed Translation Entry Areas (TEAs). DMT establishes a direct mapping from each virtual page in a Virtual Memory Area (VMA) to the corresponding PTE in a TEA. Since processes manage memory with a handful of major VMAs, the mapping can be maintained per VMA and effectively stored in a few dedicated registers. DMT further optimizes virtualized memory translation via guest-host cooperation by directly allocating guest TEAs in physical memory, bypassing intermediate virtualization layers. DMT is inherently scalable---it takes one, two, and three memory references in native, virtualized, and nested virtualized setups. Its scalability enables hardware-assisted translation for nested virtualization. Our evaluation shows that DMT significantly speeds up page walks by an average of 1.58x (1.65x with THP) in a virtualized setup, resulting in 1.20x (1.14x with THP) speedup of application execution on average. Jiyuan Zhang 0003, Weiwei Jia 0001, Siyuan Chai 0001, Peizhe Liu, Jongyul Kim 0001, Tianyin Xu |
ASPLOS (2) | 2 |
| 2024 | Effective Huge Page Strategies for TLB Miss Reduction in Nested VirtualizationabstractHuge page strategies, such as Linux Transparent Huge Page (THP), have become a prevalent solution to mitigate the performance bottleneck caused by increasingly high memory address translation overhead. However, in cloud environments, virtualization presents a two-fold challenge, exacerbating address translation overhead and undermining the effectiveness of huge page strategies. To effectively reduce address translation overhead, huge page strategies in the host and guest virtual machines (VMs) must work in concert for “proper huge page alignment”, i.e., huge pages in guest VMs being backed by host huge pages. This requires a cross-layer coordinating mechanism, which has been designed targeting non-nested virtualization settings. The paper introduces XGEMINI as an efficient solution targeting nested virtualization settings, where addressing these issues is particularly challenging, given the additional obstacles in creating synergy between host and guest VMs, due to an extra layer of page mappings by guest hypervisors. XGEMINI addresses these challenges by improving the shadow paging mechanism. Evaluation based on the KVM/Linux prototype implementation and diverse real-world applications shows XGEMINI greatly reduces TLB misses and enhances application performance in nested virtualization. Weiwei Jia 0001, Jiyuan Zhang 0003, Jianchen Shan, Xiaoning Ding |
IEEE Trans. Computers | 1 |
| 2023 | HugeGPT: Storing Guest Page Tables on Host Huge Pages to Accelerate Address TranslationabstractExpensive page table walks triggered by frequent TLB misses have incurred major performance bottlenecks for data-intensive workloads that are dominated by memory accesses with weak locality. Since it is hard to reduce TLB misses for such workloads, reducing page table walk overhead (i.e., the overhead of each TLB miss) is an increasingly important direction for improving application performance. The direction is more compelling for workloads running in virtual machines (VMs). In virtualized environments, each TLB miss triggers a two-dimensional page table walk, which has a significantly higher overhead than that on native systems. This paper presents HugeGPT, a software approach to reducing two-dimensional page table walk overhead in virtualized environments. HugeGPT ensures that page tables used in guest systems are physically held in the huge pages formed in the host. This brings two-fold benefits: 1) the number of steps walking down the host page table is reduced; 2) the misses of page walk caches incurred by accessing the leaf nodes on host page tables can be eliminated. Extensive evaluation based on the prototype implementation and diverse real-world applications shows that HugeGPT can efficiently reduce address translation overhead and improve application performance in virtualized clouds. Weiwei Jia 0001, Jiyuan Zhang 0003, Jianchen Shan, Yiming Du, Xiaoning Ding, Tianyin Xu |
PACT | 1 |
| 2023 | Making Dynamic Page Coalescing Effective on Virtualized CloudsabstractUsing huge pages has become a mainstream method to reduce address translation overhead for big memory workloads in modern computer systems. To create huge pages, system software usually uses page coalescing methods to dynamically combine contiguous base pages. Though page coalescing methods help effectively reduce address translation overhead on native systems, as the paper shows, their effectiveness is substantially undermined on virtualized platforms. Weiwei Jia 0001, Jiyuan Zhang 0003, Jianchen Shan, Xiaoning Ding |
EuroSys | 1 |
| 2023 | Reestablishing Page Placement Mechanisms for Nested VirtualizationabstractPage placement mechanisms have long been used to reduce cache conflict misses. They become more important in clouds where the emerging way-based cache partitioning is used for better workload isolation but at a cost of increased cache conflicts. However, page placement mechanisms become ineffective in virtualized environments, such as clouds, because the real locations of memory pages (i.e., their host physical addresses) are hidden from guest OSs. The paper proposesxPlaceas a solution to reestablish page placement mechanisms under the nested virtualization configuration. To keep high portability and low overhead,xPlacefollows an approach that creates a synergy between the host and guest VMs, such that the page placement mechanism inside each guest VM becomes effective even if its page placement decisions are made based on the guest physical addresses of memory pages. The paper addresses the technical issues for implementing this approach in the nested virtualization setting, particularly how to create the synergy with the obstacle created by guest hypervisors sitting between the host and guest VMs. Evaluation based on the prototype implementation and diverse real world applications shows thatxPlacecan greatly reduce cache conflicts and improve application performance in the nested environment. Xiaowei Shang, Weiwei Jia 0001, Jianchen Shan, Xiaoning Ding, Cristian Borcea |
IEEE Trans. Cloud Comput. | 2 |
| 2022 | Achieving low latency in public edges by hiding workloads mutual interferenceabstractOn multi-tenant platforms, such as public clouds and edges, workloads interfere with each other through shared resources. The performance degradation caused by such interference is a notoriously challenging problem. Though many solutions have been proposed for clouds, they can hardly help the application in edges, where workloads are mostly latency-critical, highly dynamic, and more sensitive to interference. Aggressive resource over-provisioning looks to be the only practical solution, albeit it causes significant resource waste. Weiwei Jia 0001, Jiyuan Zhang 0003, Jianchen Shan, Jing Li 0025, Xiaoning Ding |
SoCC | 1 |
| 2021 | CoPlace: Effectively Mitigating Cache Conflicts in Modern CloudsabstractSubstantial renovations in hardware cache have been focused on reducing cache interference between workloads recently. However, cache conflicts within each workload are surprisingly overlooked. The paper identifies that cache conflicts cannot be effectively reduced in virtualized clouds. Enhancements for cache partitioning, such as Intel cache allocation technology, make cache conflicts even more serious for cloud workloads. The paper proposes CoPlace as a low overhead and highly portable solution for virtualized clouds. CoPlace enhances the page placement mechanisms implemented in the host OS, such that it can collaborate with the guest OS to reduce cache conflicts. With CoPlace, the guest OS makes page placement decisions; and the host OS helps enforce the decisions. Evaluation based on the prototype implementation in Linux and KVM and diverse real world applications shows that CoPlace can significantly reduce cache conflicts and improve application performance. Xiaowei Shang, Weiwei Jia 0001, Jianchen Shan, Xiaoning Ding |
PACT | 2 |
| 2021 | Diagnosing the Interference on CPU-GPU Synchronization Caused by CPU Sharing in Multi-Tenant GPU CloudsabstractThe GPU-accelerated cloud, enabled by maturing GPU virtualization techniques, has become the most attractive platform for high-performance computing and machine learning workloads. However, it is notoriously challenging to build the multi-tenant GPU cloud where resources, like CPUs and GPUs, can be shared. One well-known and heavily studied reason is that workloads suffer from poor performance isolation and low GPU utilization when GPUs are shared. But little attention has been paid to another fundamental yet under studied problem: how sharing CPUs among GPU instances could affect the workload performance?Targeting this problem, the paper conducts experiments to measure the performance slowdown and vGPU utilization decrease under interference from CPU sharing. The results show that GPU workloads suffer from poor and unpredictable performance and heavy vGPU under-utilization because of CPU sharing. We find that such interference is the result of the complex interplay between the characteristics of CPU-GPU interactions and the special behavior of shared vCPUs: vCPU discontinuity. To diagnose how vCPU discontinuity causes the interference, the paper leverages NVIDIA Nsight Systems for fine-grained profiling and has the following findings: 1) vCPU discontinuity causes inefficient CPU-GPU synchronizations; 2) vCPU discontinuity delays task offloading to the vGPU; 3) Polling-based CPU-GPU synchronization suffers from interference more than blocking-based CPU-GPU synchronization; 4) GPU workloads with frequent task offloads and synchronizations are more vulnerable. Based on the findings, the paper proposes a novel polling-then-blocking CPU-GPU synchronization primitive. Evaluation shows that it can improve the performance by 4.2x. Youssef Elmougy, Weiwei Jia 0001, Xiaoning Ding, Jianchen Shan |
IPCCC | 2 |
| 2021 | IceClave: A Trusted Execution Environment for In-Storage ComputingabstractIn-storage computing with modern solid-state drives (SSDs) enables developers to offload programs from the host to the SSD. It has been proven to be an effective approach to alleviate the I/O bottleneck. To facilitate in-storage computing, many frameworks have been proposed. However, few of them treat the in-storage security as the first citizen. Specifically, since modern SSD controllers do not have a trusted execution environment, an offloaded (malicious) program could steal, modify, and even destroy the data stored in the SSD. Luyi Kang, Yuqi Xue, Weiwei Jia 0001, Xiaohao Wang, Jongryool Kim, Changhwan Youn, Myeong Joon Kang, Hyung Jin Lim, Bruce L. Jacob, Jian Huang 0006 |
MICRO | 3 |
| 2020 | vSMT-IO: Improving I/O Performance and Efficiency on SMT Processors in Virtualized Clouds
Weiwei Jia 0001, Jianchen Shan, Tsz On Li, Xiaowei Shang, Heming Cui, Xiaoning Ding |
USENIX ATC | 1 |
| 2018 | PLOVER: Fast, Multi-core Scalable Virtual Machine Fault-tolerance
Cheng Wang 0021, Xusheng Chen, Weiwei Jia 0001, Boxuan Li, Haoran Qiu, Shixiong Zhao, Heming Cui |
NSDI | 3 |
| 2018 | Effectively Mitigating I/O Inactivity in vCPU Scheduling
Weiwei Jia 0001, Cheng Wang 0021, Xusheng Chen, Jianchen Shan, Xiaowei Shang, Heming Cui, Xiaoning Ding, Luwei Cheng, Francis C. M. Lau 0001, Yuangang Wang |
USENIX ATC | 1 |
| 2017 | Rethinking Multicore Application Scalability on Big Virtual MachinesabstractVirtual machine (VM) sizes keep increasing in the cloud. However, little attention has been paid to analyze and understand the scalability of multicore applications on big VMs with multiple virtual CPUs (VCPUs), assuming that application scalability on VMs can be analyzed in the same ways as that on physical machines (PMs). The paper demonstrates that, since hardware CPU resource is dynamically allocated to VCPUs, the executions of multicore applications on VMs show different scalability from those on PMs. The paper systematically studies how the virtualization of CPU resource changes execution scalability, identifies key application features and system factors that affect execution scalability on VMs, and investigates possible directions to improve scalability. The paper presents a few important findings. First, the execution scalability of applications on VMs is determined by different factors than those on PMs. Second, virtualization and resource sharing can improve scalability by nature. Thus, applications may show better scalability on VMs than on PMs. Linear scalability can be achieved even when there is substantial sequential computation. Third, there is still much space to further improve execution scalability by enhancing system designs. Better scalability can be achieved by increasing allocation period length and/or matching resource allocation and workload distribution. Jianchen Shan, Weiwei Jia 0001, Xiaoning Ding |
ICPADS | 2 |