VLDB 2026 Research / reviewers in the wild / expert
Zhenlin Wang 0003
dblp:88/5294-3
· DBLP profile ↗
71ranked-venue papers
1as first author
26since 2021 · last 2026
0000-0002-0429-4371ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 51 · 1 first-author · 20 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5Artificial intelligence and machine learning · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Security and privacy · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlowGPU: Transparent and Efficient GPU Checkpointing and Restore
Zehua Yang, Yonghao Zou, Junyang Zhang 0003, Zhisheng Ye 0002, Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003, Diyu Zhou |
Euro-Par (2) | 9 |
| 2026 | Latency-SLO-Aware Memory Offloading for Large Language Model InferenceabstractOffloading large language models (LLMs) states to host memory during inference promises to reduce operational costs by supporting larger models, longer prompts, and larger batch sizes. However, the design of existing memory offloading mechanisms does not take latency service-level objectives (SLOs) into consideration. As a result, they either lead to frequent SLO violations or underutilize host memory, thereby incurring economic loss and thus defeating the purpose of memory offloading. Chenxiang Ma, Zhisheng Ye 0002, Zehua Yang, Tianhao Fu, Jiaxun Han, Jie Zhang 0048, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Yong Li 0045, Diyu Zhou |
ICS | 10 |
| 2026 | Marshaled Learning: Bridging Large Neural Networks with Memory-Constrained Trusted Execution Environments in Federated LearningabstractDespite the privacy-oriented design, federated learning (FL) remains vulnerable to privacy breaches due to the exposure of model update snapshots throughout training. Trusted Execution Environments (TEEs) offer hardware-based isolation to safeguard data and computations, providing a compelling foundation for privacy-preserving FL. However, the limited memory available in mainstream TEEs hinders the deployment of large-scale neural networks, such as GPT models, within these secure enclaves. To address this limitation, we propose Marshaled Learning, a novel FL framework that enables large neural network training across memory-constrained TEEs while ensuring strong privacy guarantees for both data and model owners. To achieve this, Marshaled Learning partitions a model into subnets and distributes them across clients according to their memory capacities, coordinating forward and backward passes across TEE-isolated environments. To mitigate the impact of heterogeneous data distributions and straggler clients, we introduce a dynamic knowledge propagation mechanism that facilitates cross-client learning and accelerates convergence. We present both theoretical convergence guarantees and empirical evaluations, demonstrating that Marshaled Learning outperforms existing FL methods by around 2% to 5% accuracy with much faster convergence rates. We also implement Marshaled Learning on commercial Azure Confidential VMs to prove its feasibility and show that it incurs only a 1 ~ 3× computational overhead compared to non-TEE settings, validating its practicality in real-world deployments. Shiwei Ding, Xiaoyong Yuan, Zhenlin Wang 0003, Lan Zhang 0005, Giuseppe Ateniese |
WACV | 3 |
| 2026 | A Comprehensive Study on Solving Memory Bloat Under VirtualizationabstractHuge pages are effective in reducing address translation overhead under virtualization. However, huge pages can lead to the memory bloat problem, which manifests in two primary forms: hot bloat and usage bloat . Hot bloat occurs when accesses to a huge page are heavily skewed towards a small subset of base pages, leading the hypervisor to (mistakenly) classify the entire huge page as hot. Hot bloat undermines several critical virtualization techniques, including tiered memory and page sharing. Usage bloatrefers to the base pages within a huge page that has not yet been allocated, causing virtual machines (VMs) to demand excessive memory. Prior work addressing memory bloat either requires hardware modification or targets a specific scenario and is not applicable to a hypervisor. This article presents HugeScope , a lightweight, effective and generic system that addresses the memory bloat problem under virtualization based on commodity hardware. HugeScope includes an efficient and precise page tracking mechanism, leveraging the other level of indirect memory translation in the hypervisor. HugeScope provides a generic framework to support page splitting and coalescing policies, considering the memory pressure, as well as the recency, frequency, and skewness of page access. Moreover, HugeScope is general and modular. It can not only be easily applied to various scenarios concerning hot bloat , including tiered memory management ( HS-TMM ) and page sharing ( HS-Share ), but also seamlessly expose its capabilities to VMs to address the usage bloat problem ( HS-HP ). Evaluation shows that HugeScope incurs less than 4% overhead, by addressing hot bloat , HS-TMM improves performance by up to 61% over vTMM while HS-Share saves 41% more memory than Ingens while offering comparable performance, and By addressing usage bloat , HS-HP can eliminate excessive memory usage, and achieve performance improvements of up to 11% over HawkEye. Chuandong Li 0004, Dong Liu 0042, Zhihong Xue, Xiaolin Wang 0001, Zhenlin Wang 0003, Yingwei Luo, Diyu Zhou |
ACM Trans. Comput. Syst. | 6 |
| 2025 | ANG: Accelerating NFA processing on GPUs via Exploring Multi-Level Fine-Grained ParallelismabstractFinite Automata (FA) processing is a core computation in various real-world applications. Over the past decades, extensive efforts have been dedicated to accelerating FA processing on modern parallel platforms, particularly GPUs, due to their high memory bandwidth and massive hardware parallelism. As Non-deterministic Finite Automata (NFA)-based applications have strong and growing demands for real-time data analytics nowadays, reducing latency in automata processing has become a critical priority. However, existing approaches face significant challenges when limited parallelism is exposed in NFA computations. In this work, we explore opportunities of introducing fine-grained parallelism from various sources and addressing the limitations of fast NFA processing. Specifically, by analyzing different NFA parallelization schemes, we identify the major performance issue caused by insufficient state-level parallelism in conventional designs. To overcome the bottleneck, this work introduces speculative parallelization tailored for GPU-based NFA processing, thus effectively exploiting fine-grained parallelism across multilevels, with a particular focus on input-chunk-level parallelism. To realize speculative parallelization in practice, we develop $A N G$, a latency-oriented NFA processing framework that overcomes key implementation challenges on GPUs. We evaluate the efficiency of ANG on a set of representative NFAs with diverse properties. Experimental results demonstrate that ANG achieves significant performance improvement compared to state-of-theart techniques, with reaching $11.74 \times$ speedup on average (and up to $49.88 \times$ in extreme cases). Yuguang Wang 0005, Yunmo Zhang, Junqiao Qiu, Zhenlin Wang 0003 |
PACT | 5 |
| 2025 | SPDK+: Low Latency or High Power Efficiency? We Take BothabstractSPDK, as one of the most efficient I/O storage software, is capable of delivering the lowest I/O latency. Unfortunately, the polling mechanism in SPDK wastes tremendous CPU clock cycles, especially under small I/O operations and low queue depths. Although SPDK supports the conventional interrupt method, it does not improve power efficiency under such circumstances. To address this issue, we propose SPDK+, which enables the user interrupt feature in the SPDK to achieve both low latency and high power efficiency. Specifically, SPDK+ employs user interrupt handling to directly process MSI-X interrupts from SSD devices and utilizes user wait instructions during IO wait periods to conserve power. The comprehensive evaluation results show that SPDK+ achieves up to 49.5% power efficiency improvement while keeping the I/O latency almost unchanged compared with SPDK. Endian Li, Shushu Yi, Qiao Li 0001, Diyu Zhou, Zhenlin Wang 0003, Xiaolin Wang 0001, Bo Mao 0003, Yingwei Luo, Ke Zhou 0001, Jie Zhang 0048 |
HotStorage | 6 |
| 2025 | Criticality-Aware Instruction-Centric Bandwidth Partitioning for Data Center ApplicationsabstractTo reduce operational costs, modern data centers co-locate high-priority latency-critical (LC) tasks and low-priority best-effort (BE) tasks on the same physical node to increase resource utilization. However, such co-location leads to contention for memory bandwidth, resulting in priority inversion, where BE tasks severely slow down LC tasks. This priority inversion often leads to violations of the quality of service (QoS) requirements for LC tasks, defeating the purpose of co-location. Prior approaches to this issue either fail to enforce the QoS requirements for LC tasks or underutilize memory bandwidth.We present Pivot, a novel bandwidth partitioning system that overcomes the limitations of prior approaches based on two key insights. First, memory accesses from LC tasks must be prioritized across all the components on the memory path rather than a single component, as done in prior work. Second, only the scheduling of a selective portion of performance-critical loads (i.e., those causing a long stall on the re-order buffer), instead of all memory accesses from LC tasks, should be prioritized. To leverage these insights, Pivot overcomes the key challenge of accurately identifying performance-critical loads while incurring minimal runtime overhead by proposing a two-phase profiling technique. Our extensive evaluation shows that Pivot improves effective machine utilization by up to $\mathbf{3 4. 5 \%}$ while increasing the throughput of the BE applications by up to $2.76 \times$ compared to state-of-the-art approaches. Liren Zhu, Liujia Li, Jie Zhang 0048, Zhenlin Wang 0003, Xiaolin Wang 0001, Yingwei Luo, Diyu Zhou |
HPCA | 7 |
| 2025 | You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation ModelsabstractFine-tuning plays a crucial role in adapting models to downstream tasks with minimal training efforts. However, the rapidly increasing size of foundation models poses a daunting challenge for accommodating foundation model fine-tuning in most commercial devices, which often have limited memory bandwidth. Techniques like model sharding and tensor parallelism address this issue by distributing computation across multiple devices to meet memory requirements. Nevertheless, these methods do not fully leverage their foundation nature in facilitating the fine-tuning process, resulting in high computational costs and imbalanced workloads. We introduce a novel Distributed Dynamic Fine-Tuning (D2FT) framework that strategically orchestrates operations across attention modules based on our observation that not all attention modules are necessary for forward and backward propagation in fine-tuning foundation models. Through three innovative selection strategies, D2FT significantly reduces the computational workload required for fine-tuning foundation models. Furthermore, D2FT addresses workload imbalances in distributed computing environments by optimizing these selection strategies via multiple knapsack optimization. Our experimental results demonstrate that the proposed D2FT framework reduces the training computational costs by 40% and training communication costs by 50% with only 1% to 2% accuracy drops on the CIFAR-10, CIFAR-100, and Stanford Cars datasets. Moreover, the results show that D2FT can be effectively extended to recent LoRA, a state-of-the-art parameter-efficient fine-tuning technique. By reducing 40% computational cost or 50% communication cost, D2FT LoRA top-1 accuracy only drops 4% to 6% on Stanford Cars dataset. The extended version of this paper can be found in http://arxiv.org/abs/2504.12471. Shiwei Ding, Lan Zhang 0005, Zhenlin Wang 0003, Giuseppe Ateniese, Xiaoyong Yuan |
IJCNN | 3 |
| 2025 | Aeolia: A Fast and Secure Userspace Interrupt-Based Storage StackabstractPolling-based userspace storage stacks achieve great I/O performance. However, they cannot efficiently and securely share disks and CPUs among multiple tasks. In contrast, interrupt-based kernel stacks inherently suffer from subpar I/O performance but achieve advantages in resource sharing. Chuandong Li 0004, Ran Yi 0004, Zonghao Zhang, Jing Liu 0074, Changwoo Min, Jie Zhang 0048, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Diyu Zhou |
SOSP | 9 |
| 2025 | CortenMM: Efficient Memory Management with Strong Correctness GuaranteesabstractModern memory management systems suffer from poor performance and subtle concurrency bugs, slowing down applications while introducing security vulnerabilities. We observe that both issues stem from the conventional design of memory management systems with two levels of abstraction: a software-level abstraction (e.g., VMA trees in Linux) and a hardware-level abstraction (typically, page tables). This design increases portability but requires correctly and efficiently synchronizing two drastically different and complex data structures, which is generally challenging. Junyang Zhang 0003, Xiangcan Xu, Yonghao Zou, Xinyi Wan 0001, Siyuan Wang 0026, Di Wang 0017, Hao Chen 0023, Lin Huang 0005, Shoumeng Yan, Yuval Tamir, Yingwei Luo, Xiaolin Wang 0001, Huashan Yu, Zhenlin Wang 0003, Hongliang Tian, Diyu Zhou |
SOSP | 17 |
| 2024 | EKRM: Efficient Key-Value Retrieval Method to Reduce Data Lookup Overhead for Redis
Xiaolin Wang 0001, Diyu Zhou, Liujia Li, Liren Zhu, Zhenlin Wang 0003, Yingwei Luo |
Euro-Par (1) | 7 |
| 2024 | Taming Hot Bloat Under Virtualization with HUGESCOPE
Chuandong Li 0004, Sai Sha, Yangqing Zeng, Xiran Yang, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Diyu Zhou |
USENIX ATC | 7 |
| 2024 | Enabling per-file data recovery from ransomware attacks via file system forensics and flash translation layer data extractionabstractAbstract Ransomware attacks are increasingly prevalent in recent years. Crypto-ransomware corrupts files on an infected device and demands a ransom to recover them. In computing devices using flash memory storage (e.g., SSD, MicroSD, etc.), existing designs recover the compromised data by extracting the entire raw flash memory image, restoring the entire external storage to a good prior state. This is feasible through taking advantage of the out-of-place updates feature implemented in the flash translation layer (FTL). However, due to the lack of “file” semantics in the FTL, such a solution does not allow a fine-grained data recovery in terms of files. Considering the file-centric nature of ransomware attacks, recovering the entire disk is mostly unnecessary. In particular, the user may just wish a speedy recovery of certain critical files after a ransomware attack. In this work, we have designed $$\textsf{FFRecovery}$$ FFRecovery , a new ransomware defense strategy that can support fine-grained per file data recovery after the ransomware attack. Our key idea is that, to restore a file corrupted by the ransomware, we (1) restore its file system metadata via file system forensics, and (2) extract its file data via raw data extraction from the FTL, and (3) assemble the corresponding file system metadata and the file data. Another essential aspect of $$\textsf{FFRecovery}$$ FFRecovery is that we add a garbage collection delay and freeze mechanism into the FTL so that no raw data will be lost prior to the recovery and, additionally, the raw data needed for the recovery can be always located. A prototype of $$\textsf{FFRecovery}$$ FFRecovery has been developed and our experiments using real-world ransomware samples demonstrate the effectiveness of $$\textsf{FFRecovery}$$ FFRecovery . We also demonstrate that $$\textsf{FFRecovery}$$ FFRecovery has negligible storage cost and performance impact. Josh Dafoe, Niusen Chen, Bo Chen 0028, Zhenlin Wang 0003 |
Cybersecur. | 4 |
| 2024 | Hardware-Software Collaborative Tiered-Memory Management Framework for VirtualizationabstractThe tiered-memory system can effectively expand the memory capacity for virtual machines (VMs). However, virtualization introduces new challenges specifically in enforcing performance isolation, minimizing context switching, and providing resource overcommit. None of the state-of-the-art designs consider virtualization and address these challenges; we observe that a VM with tiered memory incurs up to a 2× slowdown compared to a DRAM-only VM. We propose vTMM , a hardware-software collaborative tiered-memory management framework for virtualization. A key insight in vTMM is to leverage the unique system features in virtualization to meet the above challenges. vTMM automatically determines page hotness and migrates pages between fast and slow memory to achieve better performance. Specially, vTMM optimizes page tracking and migration based on page-modification logging (PML), a hardware-assisted virtualization mechanism, and adaptively distinguishes hot/cold pages through the page “temperature” sorting. vTMM also dynamically adjusts fast memory among multi-VMs on demand by using a memory pool. Further, vTMM tracks huge pages at regular-page granularity in hardware and splits/merges pages in software, realizing hybrid-grained page management and optimization. We implement and evaluate vTMM with single-grained page management on an Intel processor, and the hybrid-grained page management on a Sunway processor with hardware mode supporting hardware/software co-designs. Experiments show that vTMM outperforms existing tiered-memory management designs in virtualization. Sai Sha, Chuandong Li 0004, Xiaolin Wang 0001, Zhenlin Wang 0003, Yingwei Luo |
ACM Trans. Comput. Syst. | 4 |
| 2023 | vTMM: Tiered Memory Management for Virtual MachinesabstractThe memory demand of virtual machines (VMs) is increasing, while the traditional DRAM-only memory system has limited capacity and high power consumption. The tiered memory system can effectively expand the memory capacity and increase the cost efficiency. Virtualization introduces new challenges for memory tiering, specifically enforcing performance isolation, minimizing context switching, and providing resource overcommit. However, none of the state-of-the-art designs consider virtualization and thus address these challenges; we observe that a VM with tiered memory incurs up to a 2× slowdown compared to a DRAM-only VM. Sai Sha, Chuandong Li 0004, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003 |
EuroSys | 5 |
| 2023 | FLORIA: A Fast and Featherlight Approach for Predicting Cache PerformanceabstractThe cache Miss Ratio Curve (MRC) serves a variety of purposes such as cache partitioning, application profiling and code tuning. In this work, we propose a new metric, called cache miss distribution, that describes cache miss behavior over cache sets, for predicting cache MRCs. Based on this metric, we present FLORIA, a software-based, online approach that approximates cache MRCs on commodity systems. By polluting a tunable number of cache lines in some selected cache sets using our designed microbenchmark, the cache miss distribution for the target workload is obtained via hardware performance counters with the support of precise event based sampling (PEBS). A model is developed to predict the MRC of the target workload based on its cache miss distribution. Jun Xiao 0009, Yaocheng Xiang, Xiaolin Wang 0001, Yingwei Luo, Andy D. Pimentel, Zhenlin Wang 0003 |
ICS | 6 |
| 2023 | Multi-Tenant In-Memory Key-Value Cache Partitioning Using Efficient Random Sampling-Based LRU ModelabstractIn-memory key-value caches are widely used as a performance-critical layer in web applications, disk-based storage, and distributed systems. The Least Recently Used (LRU) replacement policy has become thede factostandard in those systems since it exploits workload locality well. However, the LRU implementation can be costly due to the rigid data structure in maintaining object priority, as well as the locks for object order updating. Redis as one of the most effective and prevalent deployed commercial systems adopts an approximated LRU policy, where the least recently used item from a small, randomly sampled set of items is chosen to evict. This random sampling-based policy is lightweight and shows its flexibility. We observe that there can exist a significant miss ratio gap between exact LRU and random sampling-based LRU under different sampling size$K$s. Therefore existing LRU miss ratio curve (MRC) construction techniques cannot be directly applied without loss of accuracy. In this article, we introduce a new probabilistic stack algorithm namedKRRto accurately model random sampling based-LRU, and extend it to handle both fixed and variable objects in key-value caches. We present an efficient stack update algorithm that reduces the expected running time of KRR significantly. To improve the performance of the in-memory multi-tenant key-value cache that utilizes random sampling-based replacement, we propose kRedis, a reference locality- and latency-aware memory partitioning scheme. kRedis guides the memory allocation among the tenants and dynamically customizes$K$to better exploit the locality of each individual tenant. Evaluation results over diverse workloads show that our model generates accurate miss ratio curves for both fixed and variable object size workloads, and enables practical, low-overhead online MRC prediction. Equipped with KRR, kRedis delivers up to a 50.2% average access latency reduction, and up to a 262.8% throughput improvement compared to Redis. Furthermore, by comparing with pRedis, a state-of-the-art design of memory allocation in Redis, kRedis shows up to 24.8% and 61.8% improvements in average access latency and throughput, respectively. Junyao Yang, Zhenlin Wang 0003 |
IEEE Trans. Cloud Comput. | 3 |
| 2022 | Tear Up the Bubble Boom: Lessons Learned From a Deep Learning Research and Development ClusterabstractWith the proliferation of deep learning, there exists a strong need to efficiently operate GPU clusters for deep learning production in giant AI companies, as well as for research and development (R&D) in small-sized research institutes and universities. Existing works have performed thorough trace analysis on large-scale production-level clusters in giant companies, which discloses the characteristics of deep learning production jobs and motivates the design of scheduling frameworks. However, R&D clusters significantly differ from production-level clusters in both job properties and user behaviors, calling for a different scheduling mechanism. In this paper, we present a detailed workload characterization of an R&D cluster, CloudBrain-I, in a research institute, Peng Cheng Laboratory. After analyzing the fine-grained resource utilization, we discover a severe problem for R&D clusters, resource underutilization, which is especially important in R&D clusters while not characterised by existing works. We further investigate two specific underutilization phenomena and conclude several implications and lessons on R&D cluster scheduling. The traces will be open-sourced to motivate further studies in the community. Zehua Yang, Zhisheng Ye 0002, Tianhao Fu, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Tianwei Zhang 0004 |
ICCD | 8 |
| 2022 | GSpecPal: Speculation-Centric Finite State Machine Parallelization on GPUsabstractFinite State Machine (FSM) plays a critical role in many real-world applications, ranging from pattern matching to network security. In recent years, significant research efforts have been made to accelerate FSM computations on different parallel platforms, including multicores, GPUs, and DRAM-based accelerators. A popular direction is the speculation-centric parallelization. Despite their abundance and promising results, the benefits of speculation-centric FSM parallelization on GPUs heavily depend on high speculation accuracy and are greatly limited by the inefficient sequential recovery. Inspired by speculative data forwarding used in Thread Level Speculation (TLS), this work addresses the existing bottlenecks by introducing speculative recovery with two heuristics for thread scheduling, which can effectively remove redundant computations and increase the GPU thread utilization. To maximize the performance of running FSMs on GPUs, this work integrates different speculative parallelization schemes into a latency-sensitive framework, GSpecPal, along with a scheme selector which aims to automatically configure the optimal GPU-based parallelization for a given FSM. Evaluation on a set of real-world FSMs with diverse characteristics confirms the effectiveness of GSpecPal. Experimental results show that GSpecPal can obtain 7.2× speedup on average (up to 20×) over the state-of-the-art on an Nvidia GeForce RTX 3090 GPU. Yuguang Wang 0005, Robbie Watling, Junqiao Qiu, Zhenlin Wang 0003 |
IPDPS | 4 |
| 2022 | Exploring GNN based program embedding technologies for binary related tasksabstractWith the rapid growth of program scale, program analysis, maintenance and optimization become increasingly diverse and complex. Applying learning-assisted methodologies onto program analysis has attracted ever-increasing attention. However, a large number of program factors including syntax structures, semantics, running platforms and compilation configurations block the effective realization of these methods. To overcome these obstacles, existing works prefer to be on a basis of source code or abstract syntax tree, but unfortunately are sub-optimal for binary-oriented analysis tasks closely related to the compilation process. To this end, we propose a new program analysis approach that aims at solving program-level and procedure-level tasks with one model, by taking advantage of the great power of graph neural networks from the level of binary code. By fusing the semantics of control flow graphs, data flow graphs and call graphs into one model, and embedding instructions and values simultaneously, our method can effectively work around emerging compilation-related problems. By testing the proposed method on two tasks, binary similarity detection and dead store prediction, the results show that our method is able to achieve as high accuracy as 83.25%, and 82.77%. Pengcheng Li 0001, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003 |
ICPC | 5 |
| 2022 | Graph Neural Networks Based Memory Inefficiency Detection Using Selective SamplingabstractProduction software of data centers oftentimes suffers from unnecessary memory inefficiencies caused by inappropriate use of data structures, conservative compiler optimizations, and so forth. Nevertheless, whole-program monitoring tools often incur incredibly high overhead due to fine-grained memory access instrumentation. Consequently, the fine-grained monitoring tools are not viable for long-running, large-scale data center applications due to strict latency criteria (e.g., service-level agreement or SLA). To this end, this work presents a novel learning-aided system, namely Puffin, to identify three kinds of unnecessary memory operations including dead stores, silent loads and silent stores, by applying gated graph neural networks onto fused static and dynamic program semantics with respect to relative positional embedding. To deploy the system in large-scale data centers, this work explores a sampling-based detection infrastructure with high efficacy and negligible overhead. We evaluate Puffin upon the well-known SPEC CPU 2017 benchmark suite for four compilation options. Experimental results show that the proposed method is able to capture the three kinds of memory inefficiencies with as high accuracy as 96% and a reduced checking overhead by$5.66\times$over the state-of-the-art tool. Pengcheng Li 0001, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Xu Liu 0001 |
SC | 5 |
| 2022 | Accelerating Address Translation for Virtualization by Leveraging Hardware ModeabstractThe overhead of memory virtualization remains nontrivial. The traditional shadow paging (TSP) resorts to a shadow page table (SPT) to achieve the native page walk speed, but page table updates require hypervisor interventions. Alternatively, nested paging enables low-overhead page table updates, but utilizes the hardware MMU to perform a long-latency two-dimensional page walk. This paper proposes new memory virtualization solutions based on hardware (machine) mode—the highest CPU privilege level in some architectures like Sunway and RISC-V. A programming interface, running in hardware mode, enables software-implementation of hardware support functions. We first proposeSoftware-based Nested Paging (SNP), which extends the software MMU to perform a two-dimensional page walk in hardware mode. Second, we presentSwift Shadow Paging (SSP), which accomplishes page table synchronization by intercepting TLB flushing in hardware mode. Finally we proposeAccelerated Shadow Paging (ASP)combining SSP and SNP. ASP handles the last-level SPT page faults by walking two-dimensional page tables in hardware mode, which eliminates most hypervisor interventions. This paper systematically compares multiple memory virtualization models by analyzing their designs and evaluating their performance both on a real system and a simulator. The experiments show that the virtualization overhead of ASP is less than 4.5% for all workloads. Sai Sha, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003 |
IEEE Trans. Computers | 5 |
| 2021 | GRAPHSPY: Fused Program Semantic Embedding through Graph Neural Networks for Memory EfficiencyabstractProduction software oftentimes suffers from unnecessary memory inefficiencies caused by inappropriate use of data structures, programming abstractions, or conservative compiler optimizations. Unfortunately, existing works often adopt a whole-program fine-grained monitoring method incurring incredibly high overhead. This work proposes a learning-aided approach to identify unnecessary memory operations, by applying several prevalent graph neural network models to extract program semantics with respect to program structure, execution semantics and dynamic states. Results show that the proposed approach captures memory inefficiencies with high accuracy of 95.27% and only around 17% overhead of the state-of-the-art. Pengcheng Li 0001, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003 |
DAC | 5 |
| 2021 | Efficient Modeling of Random Sampling-Based LRUabstractThe Miss Ratio Curve (MRC) is an important metric and effective tool for caching system performance prediction and optimization. Since the Least Recently Used (LRU) replacement policy is the de facto policy for many existing caching systems, most previous studies on efficient MRC construction are predominantly focused on the LRU replacement policy. Recently, the random sampling-based replacement mechanism, as opposed to replacement relying on the rigid LRU data structure, gains more popularity due to its lightweight and flexibility. To approximate LRU, at replacement times, the system randomly selects K objects and replaces the least recently used object among the sample. Redis implements this approximated LRU policy. We observe that there can exist a significant miss ratio gap between exact LRU and random sampling-based LRU under different sampling size K; therefore existing LRU MRC construction techniques cannot be directly applied to random sampling based LRU cache without loss of accuracy. Junyao Yang, Zhenlin Wang 0003 |
ICPP | 3 |
| 2021 | Swift shadow paging (SSP): no write-protection but following TLB flushingabstractVirtualization is a key technique for supporting cloud services and memory virtualization is a major component of virtualization technology. Common memory virtualization mechanisms include shadow paging and hardware-assisted paging. The shadow paging model needs to synchronize shadow/guest page tables whenever there is a guest page table update. In the design of traditional shadow paging (TSP), the guest page table pages are write-protected so the updates can be intercepted by the hypervisor to ensure synchronization. Frequent page table updates cause lots of VM_Exits. Researchers have developed hardware-assisted paging to eliminate this overhead. However, address translation needs to walk a two-dimensional page table. This design significantly increases the overhead of page walk. Sai Sha, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003 |
VEE | 5 |
| 2021 | Penalty- and Locality-aware Memory Allocation in Redis Using Enhanced AETabstractDue to large data volume and low latency requirements of modern web services, the use of an in-memory key-value (KV) cache often becomes an inevitable choice (e.g., Redis and Memcached). The in-memory cache holds hot data, reduces request latency, and alleviates the load on background databases. Inheriting from the traditional hardware cache design, many existing KV cache systems still use recency-based cache replacement algorithms, e.g., least recently used or its approximations. However, the diversity of miss penalty distinguishes a KV cache from a hardware cache. Inadequate consideration of penalty can substantially compromise space utilization and request service time. KV accesses also demonstrate locality, which needs to be coordinated with miss penalty to guide cache management. In this article, we first discuss how to enhance the existing cache model, the Average Eviction Time model, so that it can adapt to modeling a KV cache. After that, we apply the model to Redis and propose pRedis, Penalty- and Locality-aware Memory Allocation in Redis, which synthesizes data locality and miss penalty, in a quantitative manner, to guide memory allocation and replacement in Redis. At the same time, we also explore the diurnal behavior of a KV store and exploit long-term reuse. We replace the original passive eviction mechanism with an automatic dump/load mechanism, to smooth the transition between access peaks and valleys. Our evaluation shows that pRedis effectively reduces the average and tail access latency with minimal time and space overhead. For both real-world and synthetic workloads, our approach delivers an average of 14.0%∼52.3% latency reduction over a state-of-the-art penalty-aware cache management scheme, Hyperbolic Caching (HC), and shows more quantitative predictability of performance. Moreover, we can obtain even lower average latency (1.1%∼5.5%) when dynamically switching policies between pRedis and HC. Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003 |
ACM Trans. Storage | 4 |
| 2020 | Huge Page Friendly Virtualized Memory Management
Sai Sha, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003 |
J. Comput. Sci. Technol. | 5 |
| 2019 | pRedis: Penalty and Locality Aware Memory Allocation in RedisabstractDue to large data volume and low latency requirements of modern web services, the use of in-memory key-value (KV) cache often becomes an inevitable choice (e.g. Redis and Memcached). The in-memory cache holds hot data, reduces request latency, and alleviates the load on background databases. Inheriting from the traditional hardware cache design, many existing KV cache systems still use recency-based cache replacement algorithms, e.g., LRU or its approximations. However, the diversity of miss penalty distinguishes a KV cache from a hardware cache. Inadequate consideration of penalty can substantially compromise space utilization and request service time. KV accesses also demonstrate locality, which needs to be coordinated with miss penalty to guide cache management. Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003 |
SoCC | 4 |
| 2019 | Machine Learning for Fine-Grained Hardware Prefetcher ControlabstractModern architectures provide hardware memory prefetching capabilities which can be configured at runtime. While hardware prefetching can provide substantial performance improvements for many programs, prefetching can also increase contention for shared resources such as last-level cache and memory bandwidth. In turn, this contention can degrade performance in multi-core workloads. In this paper, we model fine-grained hardware prefetcher control as a contextual bandit, and propose a framework for learning prefetcher control policies which adjust hardware prefetching usage at runtime according to workload performance behavior. We train our policies on profiling data, wherein hardware memory prefetchers are enabled or disabled randomly at regular intervals over the course of a workload's execution. The learned prefetcher control policies provide up to a 4.3% average performance improvement over a set of memory bandwidth intensive workloads. Jason Hiebel, Laura E. Brown, Zhenlin Wang 0003 |
ICPP | 3 |
| 2019 | EMBA: Efficient Memory Bandwidth Allocation to Improve Performance on Intel Commodity ProcessorabstractOn multi-core processors, contention on shared resources such as the last level cache (LLC) and memory bandwidth may cause serious performance degradation, which makes efficient resource allocation a critical issue in data centers. Intel recently introduces Memory Bandwidth Allocation (MBA) technology on its Xeon scalable processors, which makes it possible to allocate memory bandwidth in a real system. However, how to make the most of MBA to improve system performance remains an open question. In this work, (1) we formulate a quantitative relationship between a program's performance and its LLC occupancy and memory request rate on commodity processors. (2) Guided by the performance formula, we propose a heuristic bound-aware throttling algorithm to improve system performance and (3) we further develop a hierarchical clustering method to improve the algorithm's efficiency. (4) We implement these algorithms in EMBA, a low-overhead dynamic memory bandwidth scheduling system to improve performance on Intel commodity processors. The results show that, when multiple programs run simultaneously on a multi-core processor whose memory bandwidth is saturated, the programs with high memory bandwidth demand usually use bandwidth inefficiently compared with programs with medium memory bandwidth demand from the perspective of CPU performance. By slightly throttling the former's bandwidth, we can significantly improve the performance of the latter. On average, we improve system performance by 36.9% at the expense of 8.6% bandwidth utilization rate. Yaocheng Xiang, Chencheng Ye 0001, Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003 |
ICPP | 5 |
| 2018 | DCAPS: dynamic cache allocation with partial sharingabstractIn a multicore system, effective management of shared last level cache (LLC), such as hardware/software cache partitioning, has attracted significant research attention. Some eminent progress is that Intel introduced Cache Allocation Technology (CAT) to its commodity processors recently. CAT implements way partitioning and provides software interface to control cache allocation. Unfortunately, CAT can only allocate at way level, which does not scale well for a large thread or program count to serve their various performance goals effectively. This paper proposes Dynamic Cache Allocation with Partial Sharing (DCAPS), a framework that dynamically monitors and predicts a multi-programmed workload's cache demand, and reallocates LLC given a performance target. Further, DCAPS explores partial sharing of a cache partition among programs and thus practically achieves cache allocation at a finer granularity. DCAPS consists of three parts: (1) Online Practical Miss Rate Curve (OPMRC), a low-overhead software technique to predict online miss rate curves (MRCs) of individual programs of a workload; (2) a prediction model that estimates the LLC occupancy of each individual program under any CAT allocation scheme; (3) a simulated annealing algorithm that searches for a near-optimal CAT scheme given a specific performance goal. Our experimental results show that DCAPS is able to optimize for a wide range of performance targets and can scale to a large core count. Yaocheng Xiang, Xiaolin Wang 0001, Zihui Huang, Yingwei Luo, Zhenlin Wang 0003 |
EuroSys | 6 |
| 2018 | Constructing Dynamic Policies for Paging Mode SelectionabstractVirtualization technology is a key component for data center management which allows for multiple users and applications to share a single, physical machine. Modern virtual machine monitors utilize both software and hardware-assisted paging for memory virtualization, however neither paging mode is always preferable. Previous studies have shown that dynamic selection, which at runtime selects paging modes according to relevant performance metrics, can be effective in tailoring memory virtualization to program workload. However, these approaches require low-level manual analysis, or depend on prior knowledge of workload characteristics and phasing. Jason Hiebel, Laura E. Brown, Zhenlin Wang 0003 |
ICPP | 3 |
| 2018 | Get Out of the Valley: Power-Efficient Address Mapping for GPUsabstractGPU memory systems adopt a multi-dimensional hardware structure to provide the bandwidth necessary to support 100s to 1000s of concurrent threads. On the software side, GPU-compute workloads also use multi-dimensional structures to organize the threads. We observe that these structures can combine unfavorably and create significant resource imbalance in the memory subsystem - causing low performance and poor power-efficiency. The key issue is that it is highly application-dependent which memory address bits exhibit high variability. To solve this problem, we first provide an entropy analysis approach tailored for the highly concurrent memory request behavior in GPU-compute workloads. Our window-based entropy metric captures the information content of each address bit of the memory requests that are likely to co-exist in the memory system at runtime. Using this metric, we find that GPU-compute workloads exhibit entropy valleys distributed throughout the lower order address bits. This indicates that efficient GPU-address mapping schemes need to harvest entropy from broad address-bit ranges and concentrate the entropy into the bits used for channel and bank selection in the memory subsystem. This insight leads us to propose the Page Address Entropy (PAE) mapping scheme which concentrates the entropy of the row, channel and bank bits of the input address into the bank and channel bits of the output address. PAE maps straightforwardly to hardware and can be implemented with a tree of XOR-gates. PAE improves performance by 1.31X and power-efficiency by 1.25X compared to state-of-the-art permutation-based address mapping. Xia Zhao 0004, Magnus Jahre, Zhenlin Wang 0003, Xiaolin Wang 0001, Yingwei Luo, Lieven Eeckhout |
ISCA | 4 |
| 2018 | mPart: miss-ratio curve guided partitioning in key-value storesabstractWeb applications employ key-value stores to cache the data that is most commonly accessed. The cache improves an web application's performance by serving its requests from memory, avoiding fetching them from the backend database. Since the memory space is limited, maximizing the memory utilization is a key to delivering the best performance possible. This has lead to the use of multi-tenant systems, allowing applications to share cache space. In addition, application data access patterns change over time, so the system should be adaptive in its memory allocation. Daniel Byrne, Nilufer Onder, Zhenlin Wang 0003 |
ISMM | 3 |
| 2018 | Fast Miss Ratio Curve Modeling for Storage CacheabstractThe reuse distance (least recently used (LRU) stack distance) is an essential metric for performance prediction and optimization of storage cache. Over the past four decades, there have been steady improvements in the algorithmic efficiency of reuse distance measurement. This progress is accelerating in recent years, both in theory and practical implementation. In this article, we present a kinetic model of LRU cache memory, based on the average eviction time (AET) of the cached data. The AET model enables fast measurement and use of low-cost sampling. It can produce the miss ratio curve in linear time with extremely low space costs. On storage trace benchmarks, AET reduces the time and space costs compared to former techniques. Furthermore, AET is a composable model that can characterize shared cache behavior through sampling and modeling individual programs or traces. Xiameng Hu, Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003, Chen Ding 0001, Chencheng Ye 0001 |
ACM Trans. Storage | 5 |
| 2017 | POSTER: BACM: Barrier-Aware Cache Management for Irregular Memory-Intensive GPGPU WorkloadsabstractGeneral-purpose workloads running on modern graphics processing units (GPGPUs) rely on hardware-based barriers to synchronize warps within a thread block (TB). However, imbalance may exist before reaching a barrier if a GPGPU workload contains irregular memory accesses, i.e., some warps may be critical while others may not. Ideally, cache space should be reserved for the critical warps. Unfortunately, current cache management policies are unaware of the existence of barriers and critical warps, which significantly limits the performance of irregular memory-intensive GPGPU workloads.In this work, we propose Barrier-Aware Cache Management (BACM), which is built on top of two underlying policies: a greedy policy and a friendly policy. The greedy policy does not allow non-critical warps to allocate cache lines in the L1 data cache; only critical warps can. The friendly policy allows non-critical warps to allocate cache lines but only over invalid or lower-priority cache lines. Based on the L1 data cache hit rate of non-critical warps, BACM dynamically chooses between the greedy and friendly policies. By doing so, BACM reserves more cache space to accelerate critical warps, thereby improving overall performance. Experimental results show that BACM achieves an average performance improvement of 24% and 20% compared to the GTO and BAWS policies, respectively. BACM's hardware cost is limited to 96 bytes per streaming multiprocessor. Xia Zhao 0004, Zhibin Yu 0001, Zhenlin Wang 0003, Xiaolin Wang 0001, Yingwei Luo, Lieven Eeckhout |
PACT | 4 |
| 2017 | BACM: Barrier-Aware Cache Management for Irregular Memory-Intensive GPGPU WorkloadsabstractGeneral-purpose workloads running on modern graphics processing units rely on hardware-based barriers to synchronize warps within a thread block (TB). However, imbalance may exist before reaching a barrier if a GPGPU workload contains irregular memory accesses, i.e., some warps may be critical while others may not. Ideally, cache space should be reserved for the critical warps. Unfortunately, current cache management policies are unaware of the existence of barriers and critical warps, which significantly limits the performance of irregular memory-intensive GPGPU workloads. In this paper, we propose Barrier-Aware Cache Management (BACM) which is built on top of two underlying policies: a greedy policy and a friendly policy. The greedy policy does not allow non-critical warps to allocate cache lines in the L1 data cache; only critical warps can. The friendly policy allows non-critical warps to allocate cache lines but only over invalid or lower-priority cache lines. BACM dynamically chooses between the greedy and friendly policies based on the L1 data cache hit rate for the non-critical warps. By doing so, BACM reserves more cache space to accelerate critical warps, thereby improving overall performance. Experimental results show that BACM achieves an average performance improvement of 24% and 20% compared to the GTO and BAWS policies, respectively. BACM's hardware cost is limited to 96 bytes per streaming multiprocessor. Xia Zhao 0004, Zhibin Yu 0001, Zhenlin Wang 0003, Xiaolin Wang 0001, Yingwei Luo, Lieven Eeckhout |
ICCD | 4 |
| 2017 | Evaluating the impacts of hugepage on virtual machines
Xiaolin Wang 0001, Taowei Luo, Zhenlin Wang 0003, Yingwei Luo |
Sci. China Inf. Sci. | 4 |
| 2017 | Optimizing Locality-Aware Memory Management of Key-Value CachesabstractThe in-memory cache system is a performance-critical layer in today's web server architectures. Memcached is one of the most effective, representative, and prevalent among such systems. An important problem is on its memory allocation. The default design does not make the best use of the memory. It is unable to adapt when the demand changes, a problem known as slab calcification. This paper introduces locality-aware memory allocation (LAMA), which addresses the problem by first analyzing the locality of Memcached's requests and then reassigning slabs to minimize the miss ratio or the average response time. By evaluating LAMA using various industry and academic workloads, the paper shows that LAMA outperforms existing techniques in the steady-state performance, the speed of convergence, and the ability to adapt to request pattern changes, and overcome slab calcification. The new solution is close to optimal, achieving over 98 percent of the theoretical potential. Furthermore, LAMA can also be adopted in resource partitioning to guarantee quality-of-service (QoS). Xiameng Hu, Xiaolin Wang 0001, Yingwei Luo, Chen Ding 0001, Song Jiang 0001, Zhenlin Wang 0003 |
IEEE Trans. Computers | 7 |
| 2017 | Optimal Symbiosis and Fair Scheduling in Shared CacheabstractOn multi-core processors, applications are run sharing the cache. This paper presents optimization theory to co-locate applications to minimize cache interference and maximize performance. The theory precisely specifies MRC-based composition, optimization, and correctness conditions. The paper also presents a new technique called footprint symbiosis to obtain the best shared cache performance underfair CPU allocation as well as a new sampling technique which reduces the cost of locality analysis. When sampling and optimization are combined, the paper shows that it takes less than 0.1 second analysis per program to obtain a co-run that is within 1.5 percent of the best possible performance. In an exhaustive evaluation with 12,870 tests, the best prior work improves co-run performance by 56 percent on average. The new optimization improves it by another 29 percent. Without single co-run test, footprint symbiosis is able to choose co-run choices that are just 8 percent slower than the best co-run solutions found with exhaustive testing. Xiameng Hu, Xiaolin Wang 0001, Yechen Li, Yingwei Luo, Chen Ding 0001, Zhenlin Wang 0003 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2016 | Barrier-Aware Warp Scheduling for Throughput ProcessorsabstractParallel GPGPU applications rely on barrier synchronization to align thread block activity. Few prior work has studied and characterized barrier synchronization within a thread block and its impact on performance. In this paper, we find that barriers cause substantial stall cycles in barrier-intensive GPGPU applications although GPGPUs employ lightweight hardware-support barriers. To help investigate the reasons, we define the execution between two adjacent barriers of a thread block as a warp-phase. We find that the execution progress within a warp-phase varies dramatically across warps, which we call warp-phase-divergence. While warp-phase-divergence may result from execution time disparity among warps due to differences in application code or input, and/or shared resource contention, we also pinpoint that warp-phase-divergence may result from warp scheduling. Zhibin Yu 0001, Lieven Eeckhout, Vijay Janapa Reddi, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Cheng-Zhong Xu 0001 |
ICS | 7 |
| 2016 | Kinetic Modeling of Data Eviction in Cache
Xiameng Hu, Xiaolin Wang 0001, Yingwei Luo, Chen Ding 0001, Zhenlin Wang 0003 |
USENIX ATC | 6 |
| 2016 | Dynamic Memory Balancing for VirtualizationabstractAllocating memory dynamically for virtual machines (VMs) according to their demands provides significant benefits as well as great challenges. Efficient memory resource management requires knowledge of the memory demands of applications or systems at runtime. A widely proposed approach is to construct a miss ratio curve (MRC) for a VM, which not only summarizes the current working set size (WSS) of the VM but also models the relationship between its performance and the target memory allocation size. Unfortunately, the cost of monitoring and maintaining the MRC structures is nontrivial. This article first introduces a low-cost WSS tracking system with effective optimizations on data structures, as well as an efficient mechanism to decrease the frequency of monitoring. We also propose a Memory Balancer (MEB), which dynamically reallocates guest memory based on the predicted WSS. Our experimental results show that our prediction schemes yield a high accuracy of 95.2% and low overhead of 2%. Furthermore, the overall system throughput can be significantly improved with MEB, which brings a speedup up to 7.4 for two to four VMs and 4.54 for an overcommitted system with 16 VMs. Xiaolin Wang 0001, Fang Hou 0005, Yingwei Luo, Zhenlin Wang 0003 |
ACM Trans. Archit. Code Optim. | 5 |
| 2015 | Transfer Learning-Based Co-Run Scheduling for Heterogeneous DatacentersabstractToday’s data centers are designed with multi-core CPUs where multiple virtual machines (VMs) can be co-located into one physical machine or distribute multiple computing tasks onto one physical machine. The result is co-tenancy, resource sharing and competition. Modeling and predicting such co-run interference becomes crucial for job scheduling and Quality of Service assurance. Co-locating interference can be characterized into two components, sensitivity and pressure, where sensitivity characterizes how an application’s own performance is affected by a co-run application, and pressure characterizes how much contentiousness an application exerts/brings onto the memory subsystem. Previous studies show that with simple models, sensitivity and pressure can be accurately characterized for a single machine. We extend the models to consider cross-architecture sensitivity (across different machines). Wei Kuang, Laura E. Brown, Zhenlin Wang 0003 |
AAAI | 3 |
| 2015 | Modeling Cross-Architecture Co-Tenancy Performance InterferenceabstractCloud computing has become a dominant computing paradigm to provide elastic, affordable computing resources to end users. Due to the increased computing power of modern machines powered by multi/many-core computing, data centers often co-locate multiple virtual machines (VMs) into one physical machine, resulting in co-tenancy, and resource sharing and competition. Applications or VMs co-locating in one physical machine can interfere with each other despite of the promise of performance isolation through virtualization. Modelling and predicting co-run interference therefore becomes critical for data center job scheduling and QoS (Quality of Service) assurance. Co-run interference can be categorized into two metrics, sensitivity and pressure, where the former denotes how an application's performance is affected by its co-run applications, and the latter measures how it impacts the performance of its co-run applications. This paper shows that sensitivity and pressure are both application-and architecture dependent. Further, we propose a regression model that predicts an application's sensitivity and pressure across architectures with high accuracy. This regression model enables a data center scheduler to guarantee the QoS of a VM/application when it is scheduled to co-locate with another VMs/applications. Wei Kuang, Laura E. Brown, Zhenlin Wang 0003 |
CCGRID | 3 |
| 2015 | Improving TLB Performance by Increasing Hugepage RatioabstractLinux supports transparent huge page since 2.6.38. It can automatically map huge pages. But this implementation fails to adjust to page alignment in memory allocation and thus cannot use huge page in some situations. The design is not efficient. Our work aims to increase huge page allocation, so as to improve the utilization ratio of huge page and overall performance. The experimental results show that the optimization almost reaches the upper bound of huge page utilization. This software approach delivers a notable performance improvement for a few benchmarks with moderate overhead in physical memory consumption. Taowei Luo, Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003 |
CCGRID | 5 |
| 2015 | Optimal Footprint Symbiosis in Shared CacheabstractOn multicore processors, applications are run sharing the cache. This paper presents online optimization to collocate applications to minimize cache interference to maximize performance. The paper formulates the optimization problem and solution, presents a new sampling technique for locality analysis and evaluates it in an exhaustive test of 12,870 cases. For locality analysis, previous sampling was two orders of magnitude faster than full-trace analysis. The new sampling reduces the cost by another two orders of magnitude. The best prior work improves co-run performance by 56% on average. The new optimization improves it by another 29%. When sampling and optimization are combined, the paper shows that it takes less than 0.1 second analysis per program to obtain a co-run that is within 1.5% of the best possible performance. Xiaolin Wang 0001, Yechen Li, Yingwei Luo, Xiameng Hu, Jacob Brock, Chen Ding 0001, Zhenlin Wang 0003 |
CCGRID | 7 |
| 2015 | LAMA: Optimized Locality-aware Memory Allocation for Key-value Cache
Xiameng Hu, Xiaolin Wang 0001, Yechen Li, Yingwei Luo, Chen Ding 0001, Song Jiang 0001, Zhenlin Wang 0003 |
USENIX ATC | 8 |
| 2015 | Selective switching mechanism in virtual machines via support vector machines and transfer learning
Wei Kuang, Laura E. Brown, Zhenlin Wang 0003 |
Mach. Learn. | 3 |
| 2014 | Verifying micro-architecture simulators using event tracesabstractContemporary micro-architecture research inherently relies on cycle-accurate simulators to test new ideas. Typical simulator implementations involve tens of thousands of lines of high-level code. Although general software engineering verification and validation techniques can be applied, the mere complexity of simulators makes using formal techniques difficult and calls for domain-specific knowledge to be a part of the verification process. This domain-specific information includes modeling the pipeline stages and the timing behavior of instructions with respect to these stages. Hui Meen Nyew, Nilufer Onder, Soner Önder, Zhenlin Wang 0003 |
ICS | 4 |
| 2013 | A First-Order Logic Based Framework for Verifying SimulationsabstractModern science relies on simulation techniques for understanding phenomenon, exploring design options, or evaluating models. Assuring the correctness of simulators is a key problem where a multitude of solutions ranging from manual inspection to formal verification are applicable. Formal verification incorporates the rigor necessary but not all simulators are generated from formal specifications. Manual inspection is readily available but lacks the rigor and is prone to errors. In this paper, we describe an automated verification system (AVS) where the constraints that the system must adhere to are specified by the user in general purpose first-order logic. AVS translates these constraints into a verification program that scans the simulator traceand verifies that no constraints are violated. Computer microarchitecture simulations were successfully used to demonstrate the proposed approach. This paper describes the preliminary results and discusses how artificial intelligence techniques can be used to facilitate effective run-time verification of simulators. Hui Meen Nyew, Nilufer Onder, Soner Önder, Zhenlin Wang 0003 |
AAAI | 4 |
| 2013 | Towards Eliminating Memory Virtualization Overhead
Xiaolin Wang 0001, Lingmei Weng, Zhenlin Wang 0003, Yingwei Luo |
APPT | 3 |
| 2013 | Who decides migration? A migration lock mechanism for virtual machinesabstractMigration of virtual machines is an important feature for the management of a virtualized environment. Current strategies for managing migration consider more of resource scheduling and system maintenance, ignoring possible constraints on migration from virtual machines and applications. This paper introduces a migration locking mechanism and its implementation on Xen. The virtual machine migration locking mechanism provides a standard interface for virtual machine users to control migration status of their virtual machines according to their own wishes, such as security, hardware or performance demands and so on. But forced migration locking by the end users or applications could conflict with the management decisions by managers or the virtual machine monitors. How to coordinate different requirements between virtual machine users and managers needs further discussion. Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003 |
CNSM | 3 |
| 2013 | Revisiting memory management on virtualized environmentsabstractWith the evolvement of hardware, 64-bit Central Processing Units (CPUs) and 64-bit Operating Systems (OSs) have dominated the market. This article investigates the performance of virtual memory management of Virtual Machines (VMs) with a large virtual address space in 64-bit OSs, which imposes different pressure on memory virtualization than 32-bit systems. Each of the two conventional memory virtualization approaches, Shadowing Paging (SP) and Hardware-Assisted Paging (HAP), causes different overhead for different applications. Our experiments show that 64-bit applications prefer to run in a VM using SP, while 32-bit applications do not have a uniform preference between SP and HAP. In this article, we trace this inconsistency between 32-bit applications and 64-bit applications to its root cause through a systematic empirical study in Linux systems and discover that the major overhead of SP results from memory management in the 32-bit GNU C library ( glibc ). We propose enhancements to the existing memory management algorithms, which substantially reduce the overhead of SP. Based on the evaluations using SPEC CPU2006, Parsec 2.1, and cloud benchmarks, our results show that SP, with the improved memory allocators, can compete with HAP in almost all cases, in both 64-bit and 32-bit systems. We conclude that without a significant breakthrough in HAP, researchers should pay more attention to SP, which is more flexible and cost effective. Xiaolin Wang 0001, Lingmei Weng, Zhenlin Wang 0003, Yingwei Luo |
ACM Trans. Archit. Code Optim. | 3 |
| 2012 | Live Migrating the Virtual Machine Directly Accessing a Physical NICabstractThis paper proposes a dynamic physical-virtual NIC switching. A virtual machine directly accessing the physical NIC can switch to use the virtual NIC if it is to be migrated. This solution is implemented in network configuration level in Guest OS drivers, so it can be used in different VMMs and with different NICs. The experiments illustrate that our solution does not interrupt the network connections from clients' perspective. After the virtual machine switches to use the virtual NIC, it can be migrated and the switching will not affect the migration performance, including the downtime during the migration. Xiaolin Wang 0001, Yingwei Luo, Xiaoming Li 0001, Zhenlin Wang 0003 |
APSCC | 5 |
| 2012 | A Dynamic Cache Partitioning Mechanism under Virtualization EnvironmentabstractCache sharing among multiple computing units on chip is common in today's multi-core processors, and a lot of research has focused on the effective management of shared cache. A software management method called page coloring is commonly used to divide the cache among different applications competing for the same cache entries. Both static and dynamic cache partition mechanism have been implemented in operating system or user-space level. However, few efforts have been made under virtualization environments. Our previous work has provided a static cache partition method based on page coloring in Xen, following that, a dynamic cache partition mechanism called Colored Page Migration (CoPaM) is presented in this paper. Xiaolin Wang 0001, Yechen Li, Yingwei Luo, Xiaoming Li 0001, Zhenlin Wang 0003 |
TrustCom | 6 |
| 2012 | Dynamic cache partitioning based on hot page migration
Xiaolin Wang 0001, Yechen Li, Zhenlin Wang 0003, Yingwei Luo, Xiaoming Li 0001 |
Frontiers Comput. Sci. | 4 |
| 2011 | Low Cost Working Set Size Tracking
Weiming Zhao, Xinxin Jin, Zhenlin Wang 0003, Xiaolin Wang 0001, Yingwei Luo, Xiaoming Li 0001 |
USENIX ATC | 3 |
| 2011 | Selective hardware/software memory virtualizationabstractAs virtualization becomes a key technique for supporting cloud computing, much effort has been made to reduce virtualization overhead, so a virtualized system can match its native performance. One major overhead is due to memory or page table virtualization. Conventional virtual machines rely on a shadow mechanism to manage page tables, where a shadow page table maintained by the VMM (Virtual Machine Monitor) maps virtual addresses to machine addresses while a guest maintains its own virtual to physical page table. This shadow mechanism will result in expensive VM exits whenever there is a page fault that requires synchronization between the two page tables. To avoid this cost, both Intel and AMD provide hardware assists, EPT (extended page table) and NPT (nested page table), to facilitate address translation. With the hardware assists, the MMU (Memory Management Unit) maintains an ordinary guest page table that translates virtual addresses to guest physical addresses. In addition, the extended page table as provided by EPT translates from guest physical addresses to host physical or machine addresses. NPT works in a similar style. With EPT or NPT, a guest page fault can be handled by the guest itself without triggering VM exits. However, the hardware assists do have their disadvantage compared to the conventional shadow mechanism -- the page walk yields more memory accesses and thus longer latency. Our experimental results show that neither hardware-assisted paging (HAP) nor shadow paging (SP) can be a definite winner. Despite the fact that in over half of the cases, there is no noticeable gap between the two mechanisms, an up to 34% performance gap exists for a few benchmarks. We propose a dynamic switching mechanism that monitors TLB misses and guest page faults on the fly, and dynam-ically switches between the two paging modes. Our experiments show that this new mechanism can match and, sometimes, even beat the better performance of HAP and SP. Xiaolin Wang 0001, Jiarui Zang, Zhenlin Wang 0003, Yingwei Luo, Xiaoming Li 0001 |
VEE | 3 |
| 2010 | Evaluating and Optimizing I/O Virtualization in Kernel-based Virtual Machine (KVM)
Xiaolin Wang 0001, Rongfeng Lai, Zhenlin Wang 0003, Yingwei Luo, Xiaoming Li 0001 |
NPC | 5 |
| 2010 | DMM: A dynamic memory mapping model for virtual machines
Haogang Chen 0002, Xiaolin Wang 0001, Zhenlin Wang 0003, Yingwei Luo, Xiaoming Li 0001 |
Sci. China Inf. Sci. | 3 |
| 2010 | Dynamic memory paravirtualization transparent to guest OS
Xiaolin Wang 0001, Yifeng Sun, Yingwei Luo, Zhenlin Wang 0003, Yu Li 0006, Haogang Chen 0002, Xiaoming Li 0001 |
Sci. China Inf. Sci. | 4 |
| 2009 | Fast Live Cloning of Virtual Machine Based on XenabstractVirtual Machine (VM) cloning is to create a replica of a source virtual machine (parent virtual machine); the replica, also called child virtual machine, owns exactly the same executing status as parent virtual machine. Fast live cloning guarantees that, during the period of cloning, the services running on the parent virtual machine observe no performance degradation. There are three important goals for fast live cloning: reducing the total cloning time, minimizing the suspension time of the parent virtual machine, and maximizing resource sharing between the parent virtual machine and the child virtual machine. This paper exploits Copy-on-Write (CoW) mechanism fully to minimize the time of duplicating active main memory and secondary storage of the parent virtual machine. As a result, out experiments show that the total cloning time of a virtual machine on Xen virtual machine monitor can be confined within several hundred milliseconds and the downtime of parent virtual machine is limited to tens of milliseconds, close to VMM’s scheduling interval. Experiments also show that the parent virtual machine highly shares memory with the child virtual machine as long as both virtual machines continue executing the same applications. Yifeng Sun, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Haogang Chen 0002, Xiaoming Li 0001 |
HPCC | 4 |
| 2009 | REMOCA: Hypervisor Remote Disk CacheabstractIn virtual machine (VM) systems, with the increase in the number of VMs and the demands of applications, the main memory is becoming a bottleneck of application performance. To improve paging performance for memory-intensive or I/O-intensive workloads, we propose the hypervisor REMOte disk CAche (REMOCA), which allows a virtual machine to use the memory resources on other physical machines as its cache between its virtual memory and virtual disk devices. The goal of REMOCA is to reduce disk accesses, which is much slower than transferring memory pages over modern interconnect networks. As a result, the average disk I/O latency can be improved. REMOCA is implemented within the hypervisor, by intercepting guest events such as page evictions and disk accesses. This design is transparent to the applications, and is compatible with existing techniques like ballooning and ghost buffer. Moreover, a combination of them can provide a more flexible resource management policy. Our experimental results show that REMOCA can efficiently alleviate the impact of thrashing behavior, and also significantly improve the performance for real-world I/O intensive applications. Haogang Chen 0002, Xiaolin Wang 0001, Zhenlin Wang 0003, Xinxin Jin, Yingwei Luo, Xiaoming Li 0001 |
ISPA | 3 |
| 2009 | A Simple Cache Partitioning Approach in a Virtualized EnvironmentabstractVirtualization is often used in systems for the purpose of offering isolation among applications running in separate virtual machines (VM). Current virtual machine monitors (VMMs) have done a decent job in resource isolation in memory, CPU and I/O devices. However, when looking further into the usage of lower-level shared cache, we notice that one virtual machine’s cache behavior may interfere with another’s due to the uncontrolled cache sharing. In this situation, performance isolation cannot be guaranteed. This paper presents a cache partitioning approach which can be implemented in the VMM. We have implemented this mechanism in Xen VMM using the page coloring technique traditionally applied to the OS. Our VMM-based implementation is fully transparent to the guest OSes. It thus shows the advantages of simplicity and flexibility. Our evaluation shows that our cache partitioning method can work efficiently and improve the performance of co-scheduled applications running within different VMs. In the concurrent workloads selected from the SPEC CPU 2006 benchmarks, our technique achieves a performance improvement by up to 19% for the most sensitive workloads Xinxin Jin, Haogang Chen 0002, Xiaolin Wang 0001, Zhenlin Wang 0003, Yingwei Luo, Xiaoming Li 0001 |
ISPA | 4 |
| 2009 | Dynamic memory balancing for virtual machinesabstractVirtualization essentially enables multiple operating systems and applications to run on one physical computer by multiplexing hardware resources. A key motivation for applying virtualization is to improve hardware resource utilization while maintaining reasonable quality of service. However, such a goal cannot be achieved without efficient resource management. Though most physical resources, such as processor cores and I/O devices, are shared among virtual machines using time slicing and can be scheduled flexibly based on priority, allocating an appropriate amount of main memory to virtual machines is more challenging. Different applications have different memory requirements. Even a single application shows varied working set sizes during its execution. An optimal memory management strategy under a virtualized environment thus needs to dynamically adjust memory allocation for each virtual machine, which further requires a prediction model that forecasts its host physical memory needs on the fly. This paper introduces MEmory Balancer (MEB) which dynamically monitors the memory usage of each virtual machine, accurately predicts its memory needs, and periodically reallocates host memory. MEB uses two effective memory predictors which, respectively, estimate the amount of memory available for reclaiming without a notable performance drop, and additional memory required for reducing the virtual machine paging penalty. Our experimental results show that our prediction schemes yield high accuracy and low overhead. Furthermore, the overall system throughput can be significantly improved with MEB. Copyright © 2009 ACM. Weiming Zhao, Zhenlin Wang 0003 |
VEE | 2 |
| 2008 | Live and incremental whole-system migration of virtual machines using block-bitmapabstractIn this paper, we describe a whole-system live migration scheme, which transfers the whole system run-time state, including CPU state, memory data, and local disk storage, of the virtual machine (VM). To minimize the downtime caused by migrating large disk storage data and keep data integrity and consistency, we propose a three-phase migration (TPM) algorithm. To facilitate the migration back to initial source machine, we use an incremental migration (IM) algorithm to reduce the amount of the data to be migrated. Block-bitmap is used to track all the write accesses to the local disk storage during the migration. Synchronization of the local disk storage in the migration is performed according to the block-bitmap. Experiments show that our algorithms work well even when I/O-intensive workloads are running in the migrated VM. The downtime of the migration is around 100 milliseconds, close to shared-storage migration. Total migration time is greatly reduced using IM. The block-bitmap based synchronization mechanism is simple and effective. Performance overhead of recording all the writes on migrated VM is very low. Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Yifeng Sun, Haogang Chen 0002 |
CLUSTER | 4 |
| 2006 | Path-Based Reuse Distance Analysis
Changpeng Fang, Steve Carr 0001, Soner Önder, Zhenlin Wang 0003 |
CC | 4 |
| 2006 | Feedback-directed memory disambiguation through store distance analysisabstractFeedback-directed optimization has developed into an increasingly important tool in designing optimizing compilers. Based upon profiling, memory distance analysis has shown much promise in predicting data locality and memory dependences, and has seen use in locality based optimizations and memory disambiguation. In this paper, we apply a form of memory distance, called store distance, to the problem of memory disambiguation in out-of-order issue processors. Store distance is defined as the number of store references between a load and the previous store accessing the same memory location. By generating a representative store distance for each load instruction, we can apply a compiler/micro-architecture cooperative scheme to direct run-time load speculation. Using store distance, the processor can, in most cases, accurately determine on which specific store instruction a load depends according to its store distance annotation. Our experiments show that the proposed store distance method performs much better than the previous distance based memory disambiguation scheme, and yields a performance very close to perfect memory disambiguation. The store distance based scheme also outperforms the store set technique with a relatively small predictor space and achieves performance comparable to that of a 16K-entry store set implementation for both floating point and integer programs. Changpeng Fang, Steve Carr 0001, Soner Önder, Zhenlin Wang 0003 |
ICS | 4 |
| 2004 | The garbage collection advantage: improving program localityabstractAs improvements in processor speed continue to outpace improvements in cache and memory speed, poor locality increasingly degrades performance. Because copying garbage collectors move objects, they have an opportunity to improve locality. However, no static copying order is guaranteed to match program traversal orders. This paper introduces online object reordering (OOR) which includes a new dynamic, online class analysis for Java that detects program traversal patterns and exploits them in a copying collector. OOR uses run-time method sampling that drives just-in-time (JIT) compilation. For each hot (frequently executed) method, OOR analysis identifies the hot field accesses. At garbage collection time, the OOR collector then copies referents of hot fields together with their parent. Enhancements include static analysis to exclude accesses in cold basic blocks, heuristics that decay heat to respond to phase changes, and a separate space for hot objects. The overhead of OOR is on average negligible and always less than 2% on Java benchmarks in Jikes RVM with MMTk. We compare program performance of OOR to static class-oblivious copying orders (e.g., breadth and depth first). Performance variation due to static orders is often low, but can be up to 25%. In contrast, OOR matches or improves upon the best static order since its history-based copying tunes memory layout to program traversal. Xianglong Huang, Steve Blackburn, Kathryn S. McKinley, J. Eliot B. Moss, Zhenlin Wang 0003, Perry Cheng |
OOPSLA | 5 |
| 2003 | Guided Region Prefetching: A Cooperative Hardware/Software ApproachabstractDespite large caches, main-memory access latencies still cause significant performance losses in many applications. Numerous hardware and software prefetching schemes have been proposed to tolerate these latencies. Software prefetching typically provides better prefetch accuracy than hardware, but is limited by prefetch instruction overheads and the compiler's limited ability to schedule prefetches sufficiently far in advance to cover level-two cache miss latencies. Hardware prefetching can be effective at hiding these large latencies, but generates many useless prefetches and consumes considerable memory bandwidth. We propose a cooperative hardware-software prefetching scheme called guided region prefetching (GRP), which uses compiler-generated hints encoded in load instructions to regulate an aggressive hardware prefetching engine. We compare GRP against a sophisticated pure hardware stride prefetcher and a scheduled region prefetching (SRP) engine. SRP and GRP show the best performance, with respective 22% and 21% gains over no prefetching, but SRP incurs 180% extra memory traffic-nearly tripling bandwidth requirements. GRP achieves performance close to SRP, but with a mere eighth of the extra prefetching traffic, a 23% increase over no prefetching. The GRP hardware-software collaboration thus combines the accuracy of compiler-based program analysis with the performance potential of aggressive hardware prefetching, bringing the performance gap versus a perfect L2 cache under 20%. Zhenlin Wang 0003, Doug Burger, Steven K. Reinhardt, Kathryn S. McKinley, Charles C. Weems |
ISCA | 1 |