Yingwei Luo

dblp:70/4037 · DBLP profile ↗
← Back
94ranked-venue papers
12as first author
34since 2021 · last 2026
0000-0002-7903-0717ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 47 · 1 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 4 first-authorSoftware engineering, systems software and programming languages · 14 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 5 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FlowGPU: Transparent and Efficient GPU Checkpointing and Restore
Zehua Yang, Yonghao Zou, Junyang Zhang 0003, Zhisheng Ye 0002, Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003, Diyu Zhou
Euro-Par (2)8
2026 Xerxes: Extensive Exploration of Scalable Hardware Systems with CXL-Based Simulation Framework
Yuda An, Shushu Yi, Bo Mao 0003, Qiao Li 0001, Mingzhe Zhang 0005, Diyu Zhou, Ke Zhou 0001, Nong Xiao 0001, Guangyu Sun 0003, Yingwei Luo, Jie Zhang 0048
FAST10
2026 Latency-SLO-Aware Memory Offloading for Large Language Model Inference
abstract
Offloading large language models (LLMs) states to host memory during inference promises to reduce operational costs by supporting larger models, longer prompts, and larger batch sizes. However, the design of existing memory offloading mechanisms does not take latency service-level objectives (SLOs) into consideration. As a result, they either lead to frequent SLO violations or underutilize host memory, thereby incurring economic loss and thus defeating the purpose of memory offloading.
Chenxiang Ma, Zhisheng Ye 0002, Zehua Yang, Tianhao Fu, Jiaxun Han, Jie Zhang 0048, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Yong Li 0045, Diyu Zhou
ICS8
2026 A Comprehensive Study on Solving Memory Bloat Under Virtualization
abstract
Huge pages are effective in reducing address translation overhead under virtualization. However, huge pages can lead to the memory bloat problem, which manifests in two primary forms: hot bloat and usage bloat . Hot bloat occurs when accesses to a huge page are heavily skewed towards a small subset of base pages, leading the hypervisor to (mistakenly) classify the entire huge page as hot. Hot bloat undermines several critical virtualization techniques, including tiered memory and page sharing. Usage bloatrefers to the base pages within a huge page that has not yet been allocated, causing virtual machines (VMs) to demand excessive memory. Prior work addressing memory bloat either requires hardware modification or targets a specific scenario and is not applicable to a hypervisor. This article presents HugeScope , a lightweight, effective and generic system that addresses the memory bloat problem under virtualization based on commodity hardware. HugeScope includes an efficient and precise page tracking mechanism, leveraging the other level of indirect memory translation in the hypervisor. HugeScope provides a generic framework to support page splitting and coalescing policies, considering the memory pressure, as well as the recency, frequency, and skewness of page access. Moreover, HugeScope is general and modular. It can not only be easily applied to various scenarios concerning hot bloat , including tiered memory management ( HS-TMM ) and page sharing ( HS-Share ), but also seamlessly expose its capabilities to VMs to address the usage bloat problem ( HS-HP ). Evaluation shows that HugeScope incurs less than 4% overhead, by addressing hot bloat , HS-TMM improves performance by up to 61% over vTMM while HS-Share saves 41% more memory than Ingens while offering comparable performance, and By addressing usage bloat , HS-HP can eliminate excessive memory usage, and achieve performance improvements of up to 11% over HawkEye.
Chuandong Li 0004, Dong Liu 0042, Zhihong Xue, Xiaolin Wang 0001, Zhenlin Wang 0003, Yingwei Luo, Diyu Zhou
ACM Trans. Comput. Syst.7
2025 SPDK+: Low Latency or High Power Efficiency? We Take Both
abstract
SPDK, as one of the most efficient I/O storage software, is capable of delivering the lowest I/O latency. Unfortunately, the polling mechanism in SPDK wastes tremendous CPU clock cycles, especially under small I/O operations and low queue depths. Although SPDK supports the conventional interrupt method, it does not improve power efficiency under such circumstances. To address this issue, we propose SPDK+, which enables the user interrupt feature in the SPDK to achieve both low latency and high power efficiency. Specifically, SPDK+ employs user interrupt handling to directly process MSI-X interrupts from SSD devices and utilizes user wait instructions during IO wait periods to conserve power. The comprehensive evaluation results show that SPDK+ achieves up to 49.5% power efficiency improvement while keeping the I/O latency almost unchanged compared with SPDK.
Endian Li, Shushu Yi, Qiao Li 0001, Diyu Zhou, Zhenlin Wang 0003, Xiaolin Wang 0001, Bo Mao 0003, Yingwei Luo, Ke Zhou 0001, Jie Zhang 0048
HotStorage9
2025 InstAttention: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference
abstract
The widespread of Large Language Models (LLMs) marks a significant milestone in generative AI. Nevertheless, the increasing context length and batch size in offline LLM inference escalate the memory requirement of the key-value (KV) cache, which imposes a huge burden on the GPU VRAM, especially for resource-constrained scenarios (e.g., edge computing). Several cost-effective solutions leverage host memory or SSDs to reduce storage costs for offline inference scenarios and improve the throughput. Nevertheless, they suffer from significant performance penalties imposed by intensive KV cache accesses due to limited PCIe bandwidth. To address these issues, we propose InstAttention, a novel LLM inference system that offloads the most performance-critical computation (i.e., attention in decoding phase) and data (i.e., KV cache) parts to Computational Storage Drives (CSDs), which minimize the enormous KV transfer overheads. InstAttention designs a dedicated flashaware in-storage attention engine with KV cache management mechanisms to exploit the high internal bandwidths of CSDs instead of being limited by the PCIe bandwidth. The optimized P2P transmission between GPU and CSDs further reduces data migration overheads. Experimental results demonstrate that for a 13B model using an NVIDIA A6000 GPU, InstAttention improves throughput for long-sequence inference by up to $11.1 \times$, compared to existing SSD-based solutions such as FlexGen.
Xiurui Pan, Endian Li, Qiao Li 0001, Shengwen Liang, Yizhou Shan, Ke Zhou 0001, Yingwei Luo, Xiaolin Wang 0001, Jie Zhang 0048
HPCA7
2025 Criticality-Aware Instruction-Centric Bandwidth Partitioning for Data Center Applications
abstract
To reduce operational costs, modern data centers co-locate high-priority latency-critical (LC) tasks and low-priority best-effort (BE) tasks on the same physical node to increase resource utilization. However, such co-location leads to contention for memory bandwidth, resulting in priority inversion, where BE tasks severely slow down LC tasks. This priority inversion often leads to violations of the quality of service (QoS) requirements for LC tasks, defeating the purpose of co-location. Prior approaches to this issue either fail to enforce the QoS requirements for LC tasks or underutilize memory bandwidth.We present Pivot, a novel bandwidth partitioning system that overcomes the limitations of prior approaches based on two key insights. First, memory accesses from LC tasks must be prioritized across all the components on the memory path rather than a single component, as done in prior work. Second, only the scheduling of a selective portion of performance-critical loads (i.e., those causing a long stall on the re-order buffer), instead of all memory accesses from LC tasks, should be prioritized. To leverage these insights, Pivot overcomes the key challenge of accurately identifying performance-critical loads while incurring minimal runtime overhead by proposing a two-phase profiling technique. Our extensive evaluation shows that Pivot improves effective machine utilization by up to $\mathbf{3 4. 5 \%}$ while increasing the throughput of the BE applications by up to $2.76 \times$ compared to state-of-the-art approaches.
Liren Zhu, Liujia Li, Jie Zhang 0048, Zhenlin Wang 0003, Xiaolin Wang 0001, Yingwei Luo, Diyu Zhou
HPCA9
2025 BGEM: Bipartite Graph-based Extension Modeling in Graph Neural Network for QoS Prediction
abstract
With the advancement of cloud computing and the proliferation of internet services, Quality of Service (QoS) has emerged as a critical metric for service caller, necessitating increasingly accurate predictive capabilities. While existing neural network methods have made significant progress, they have often overlooked the implicit complex interactions manifest in the distribution of QoS values. To overcome this deficiency, we introduce a Bipartite Graph-based Extension Modeling (BGEM) method with user-service interaction and a custom GCN module named Hidden Relationship Perception(HRP-GCN). In BGEM, we model the potential interactions between user-service and leverage the Jensen-Shannon divergence to quantify the similarity of QoS distribution patterns in users and services, resulting in the BGE-Graph. Subsequently, we employ a subgraph sampling algorithm to identify the optimal subgraph, and then fed it into HRP-GCN module for the prediction of unknown QoS values. Extensive experiments conducted on real-world datasets demonstrate that our proposed BGEM method outperforms current state-of-the-art approaches in QoS prediction, particularly under low-density conditions, where it still maintains superior performance.
Yingwei Luo, Yugen Du, Guoxing Tang
IJCNN1
2025 MFAE: Multi-Feature-Aware Expert Modeling for Web Service QoS Prediction (S)
abstract
With the rapid proliferation of Web services, accurately and efficiently predicting Quality of Service (QoS) has become a critical challenge in the field of service recommendation.However, existing deep learning approaches often focus on isolated feature types of either users or services and lack the capacity to comprehensively model the complex interactions between them.To address this limitation, this paper proposes a novel QoS prediction model named Multi-Feature-Aware Expert Modeling (MFAE).MFAE systematically extracts and processes five heterogeneous types of features associated with users, services, and their interactions: ID features, network topology features, geo-spatial features, similarity features, and 3-sigma-based outlier features.To effectively handle the heterogeneity among these feature types, the model constructs a network of expert groups, where each expert group consists of multiple Multi-Layer Perceptrons (MLPs) dedicated to deep representation learning for a specific feature category.The outputs from these expert groups are then fused through a weighted aggregation to generate the final QoS prediction.We evaluate MFAE on a real-world QoS dataset.Experimental results show that when the matrix density ranges from 2.5% to 10%, MFAE consistently achieves lower Mean Absolute Error (MAE) and Root Mean Square Error (RMSE) compared to six baseline methods, demonstrating its effectiveness and robustness in sparse scenarios.
Yunpeng Han, Yugen Du, Yingwei Luo, Guoxing Tang, Benchi Ma
SEKE4
2025 Aeolia: A Fast and Secure Userspace Interrupt-Based Storage Stack
abstract
Polling-based userspace storage stacks achieve great I/O performance. However, they cannot efficiently and securely share disks and CPUs among multiple tasks. In contrast, interrupt-based kernel stacks inherently suffer from subpar I/O performance but achieve advantages in resource sharing.
Chuandong Li 0004, Ran Yi 0004, Zonghao Zhang, Jing Liu 0074, Changwoo Min, Jie Zhang 0048, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Diyu Zhou
SOSP7
2025 CortenMM: Efficient Memory Management with Strong Correctness Guarantees
abstract
Modern memory management systems suffer from poor performance and subtle concurrency bugs, slowing down applications while introducing security vulnerabilities. We observe that both issues stem from the conventional design of memory management systems with two levels of abstraction: a software-level abstraction (e.g., VMA trees in Linux) and a hardware-level abstraction (typically, page tables). This design increases portability but requires correctly and efficiently synchronizing two drastically different and complex data structures, which is generally challenging.
Junyang Zhang 0003, Xiangcan Xu, Yonghao Zou, Xinyi Wan 0001, Siyuan Wang 0026, Di Wang 0017, Hao Chen 0023, Lin Huang 0005, Shoumeng Yan, Yuval Tamir, Yingwei Luo, Xiaolin Wang 0001, Huashan Yu, Zhenlin Wang 0003, Hongliang Tian, Diyu Zhou
SOSP14
2025 ASTERINAS: A Linux ABI-Compatible, Rust-Based Framekernel OS with a Small and Sound TCB
Yuke Peng, Hongliang Tian, Junyang Zhang 0003, Jinyi Xian, Xiaolin Wang 0001, Chenren Xu, Diyu Zhou, Yingwei Luo, Shoumeng Yan, Yinqian Zhang
USENIX ATC11
2025 A QoS Prediction Framework via Utility Maximization and Region-Aware Matrix Factorization
abstract
With the surge of Web services, users are more concerned about Quality-of-Service (QoS) information when choosing Web services with similar functionalities. Today, effectively and accurately predicting QoS values is a tough challenge. Typically, traditional methods only use the QoS values provided by users to predict the missing QoS values, ignoring the arbitrariness of some users in providing observed QoS values and failing to consider the existence of anomalous QoS values with contingencies caused by some unstable Web services. Taking into account the above, this article proposes HyLoReF-us, a new framework for QoS prediction. HyLoReF-us uses the user reputation to measure the trustworthiness of users and the service reputation to measure the stability of web services. First, considering the utility generated by the invocation between users and Web services, HyLoReF-us employs a Logit model to calculate the user reputation and service reputation. Second, after combining the location information of users and services, as well as their reputations, HyLoReF-us obtains QoS predictions through an improved Matrix Factorization (MF) model. Finally, a series of experiments were conducted on the standard WS-DREAM dataset. Experimental results show that HyLoReF-us outperforms current state-of-the-art or baseline methods at Matrix Densities (MD) from 5% to 30%.
Yugen Du, Guoxing Tang, Yingwei Luo, Hanting Wang
IEEE Trans. Serv. Comput.5
2024 EKRM: Efficient Key-Value Retrieval Method to Reduce Data Lookup Overhead for Redis
Xiaolin Wang 0001, Diyu Zhou, Liujia Li, Liren Zhu, Zhenlin Wang 0003, Yingwei Luo
Euro-Par (1)8
2024 HyLoReF: A Reputation Based QoS Prediction Framework using Hybrid Location Information
abstract
With the proliferation of Web services, users pay more attention to Quality of Service (QoS) information when choosing Web services with similar functionalities. Predicting QoS values effectively and accurately is a difficult challenge. In the real world, some users are strongly subjective in submitting QoS observations and some Web services suffer from instability caused by bugs. Therefore, this paper uses user reputation to measure the reliability of users and service reputation to measure the stability of Web services. We propose HyLoReF, a QoS prediction framework based on reputation and hybrid location information. HyLoReF uses Logit model to compute user reputation and service reputation, and combines them into Matrix Factorization (MF) model to improve the QoS prediction accuracy. Experimental results show that HyLoReF outperforms baseline methods and state-of-the-art models on the Web services standard dataset WS-DREAM [1] when the Matrix Density (MD) is in the interval of 5% to 30%.
Yugen Du, Hanting Wang, Yingwei Luo, Benchi Ma, Guoxing Tang
ICWS5
2024 Characterization of Large Language Model Development in the Datacenter
Qinghao Hu 0004, Zhisheng Ye 0002, Zerui Wang, Guoteng Wang, Meng Zhang 0045, Qiaoling Chen, Peng Sun 0006, Dahua Lin, Xiaolin Wang 0001, Yingwei Luo, Yonggang Wen 0001, Tianwei Zhang 0004
NSDI10
2024 RAHN: A Reputation Based Hourglass Network for Web Service QoS Prediction (S)
abstract
As the homogenization of Web services becomes more and more common, the difficulty of service recommendation is gradually increasing. How to predict Quality of Service (QoS) more efficiently and accurately becomes an important challenge for service recommendation. Considering the excellent role of reputation and deep learning (DL) techniques in the field of QoS prediction, we propose a reputation and DL based QoS prediction network, RAHN, which contains the Reputation Calculation Module (RCM), the Latent Feature Extraction Module (LFEM), and the QoS Prediction Hourglass Network (QPHN). RCM obtains the user reputation and the service reputation by using a clustering algorithm and a Logit model. LFEM extracts latent features from known information to form an initial latent feature vector. QPHN aggregates latent feature vectors with different scales by using Attention Mechanism, and can be stacked multiple times to obtain the final latent feature vector for prediction. We evaluate RAHN on a real QoS dataset. The experimental results show that the Mean Absolute Error (MAE) and Root Mean Square Error (RMSE) of RAHN are smaller than the six baseline methods.
Yugen Du, Guoxing Tang, Yingwei Luo, Benchi Ma
SEKE4
2024 Fuzzy Information Entropy and Region Biased Matrix Factorization for Web Service QoS Prediction
abstract
Nowadays, there are many similar services available on the internet, making Quality of Service (QoS) a key concern for users.Since collecting QoS values for all services through user invocations is impractical, predicting QoS values is a more feasible approach.Matrix factorization is considered an effective prediction method.However, most existing matrix factorization algorithms focus on capturing global similarities between users and services, overlooking the local similarities between users and their similar neighbors, as well as the non-interactive effects between users and services.This paper proposes a matrix factorization approach based on user information entropy and region bias, which utilizes a similarity measurement method based on fuzzy information entropy to identify similar neighbors of users.Simultaneously, it integrates the region bias between each user and service linearly into matrix factorization to capture the noninteractive features between users and services.This method demonstrates improved predictive performance in more realistic and complex network environments.Additionally, numerous experiments are conducted on real-world QoS datasets.The experimental results show that the proposed method outperforms some of the state-of-the-art methods in the field at matrix densities ranging from 5% to 20%.
Guoxing Tang, Yugen Du, Yingwei Luo, Benchi Ma
SEKE4
2024 Taming Hot Bloat Under Virtualization with HUGESCOPE
Chuandong Li 0004, Sai Sha, Yangqing Zeng, Xiran Yang, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Diyu Zhou
USENIX ATC5
2024 Hardware-Software Collaborative Tiered-Memory Management Framework for Virtualization
abstract
The tiered-memory system can effectively expand the memory capacity for virtual machines (VMs). However, virtualization introduces new challenges specifically in enforcing performance isolation, minimizing context switching, and providing resource overcommit. None of the state-of-the-art designs consider virtualization and address these challenges; we observe that a VM with tiered memory incurs up to a 2× slowdown compared to a DRAM-only VM. We propose vTMM , a hardware-software collaborative tiered-memory management framework for virtualization. A key insight in vTMM is to leverage the unique system features in virtualization to meet the above challenges. vTMM automatically determines page hotness and migrates pages between fast and slow memory to achieve better performance. Specially, vTMM optimizes page tracking and migration based on page-modification logging (PML), a hardware-assisted virtualization mechanism, and adaptively distinguishes hot/cold pages through the page “temperature” sorting. vTMM also dynamically adjusts fast memory among multi-VMs on demand by using a memory pool. Further, vTMM tracks huge pages at regular-page granularity in hardware and splits/merges pages in software, realizing hybrid-grained page management and optimization. We implement and evaluate vTMM with single-grained page management on an Intel processor, and the hybrid-grained page management on a Sunway processor with hardware mode supporting hardware/software co-designs. Experiments show that vTMM outperforms existing tiered-memory management designs in virtualization.
Sai Sha, Chuandong Li 0004, Xiaolin Wang 0001, Zhenlin Wang 0003, Yingwei Luo
ACM Trans. Comput. Syst.5
2023 vTMM: Tiered Memory Management for Virtual Machines
abstract
The memory demand of virtual machines (VMs) is increasing, while the traditional DRAM-only memory system has limited capacity and high power consumption. The tiered memory system can effectively expand the memory capacity and increase the cost efficiency. Virtualization introduces new challenges for memory tiering, specifically enforcing performance isolation, minimizing context switching, and providing resource overcommit. However, none of the state-of-the-art designs consider virtualization and thus address these challenges; we observe that a VM with tiered memory incurs up to a 2× slowdown compared to a DRAM-only VM.
Sai Sha, Chuandong Li 0004, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003
EuroSys3
2023 FLORIA: A Fast and Featherlight Approach for Predicting Cache Performance
abstract
The cache Miss Ratio Curve (MRC) serves a variety of purposes such as cache partitioning, application profiling and code tuning. In this work, we propose a new metric, called cache miss distribution, that describes cache miss behavior over cache sets, for predicting cache MRCs. Based on this metric, we present FLORIA, a software-based, online approach that approximates cache MRCs on commodity systems. By polluting a tunable number of cache lines in some selected cache sets using our designed microbenchmark, the cache miss distribution for the target workload is obtained via hardware performance counters with the support of precise event based sampling (PEBS). A model is developed to predict the MRC of the target workload based on its cache miss distribution.
Jun Xiao 0009, Yaocheng Xiang, Xiaolin Wang 0001, Yingwei Luo, Andy D. Pimentel, Zhenlin Wang 0003
ICS4
2022 Original Content Is All You Need! an Empirical Study on Leveraging Answer Summary for WikiHowQA Answer Selection Task
abstract
Answer selection task requires finding appropriate answers to questions from informative but crowdsourced candidates. A key factor impeding its solution by current answer selection approaches is the redundancy and lengthiness issues of crowdsourced answers. Recently, Deng et al. (2020) constructed a new dataset, WikiHowQA, which contains a corresponding reference summary for each original lengthy answer. And their experiments show that leveraging the answer summaries helps to attend the essential information in original lengthy answers and improve the answer selection performance under certain circumstances. However, when given a question and a set of long candidate answers, human beings could effortlessly identify the correct answer without the aid of additional answer summaries since the original answers contain all the information volume that answer summaries contain. In addition, pretrained language models have been shown superior or comparable to human beings on many natural language processing tasks. Motivated by those, we design a series of neural models, either pretraining-based or non-pretraining-based, to check wether the additional answer summaries are helpful for ranking the relevancy degrees of question-answer pairs on WikiHowQA dataset. Extensive automated experiments and hand analysis show that the additional answer summaries are not useful for achieving the best performance.
Liang Wen, Houfeng Wang, Yingwei Luo, Xiaolin Wang 0001, Xiaodong Zhang 0022, Zhicong Cheng, Dawei Yin 0001
COLING4
2022 M3: A Multi-View Fusion and Multi-Decoding Network for Multi-Document Reading Comprehension
abstract
Multi-document reading comprehension task requires collecting evidences from different documents for answering questions.Previous research works either use the extractive modeling method to naively integrate the scores from different documents on the encoder side or use the generative modeling method to collect the clues from different documents on the decoder side individually.However, any single modeling method cannot make full of the advantages of both.In this work, we propose a novel method that tries to employ a multi-view fusion and multi-decoding mechanism to achieve it.For one thing, our approach leverages question-centered fusion mechanism and cross-attention mechanism to gather finegrained fusion of evidence clues from different documents in the encoder and decoder concurrently.For another, our method simultaneously employs both the extractive decoding approach and the generative decoding method to effectively guide the training process.Compared with existing methods, our method can perform both extractive decoding and generative decoding independently and optionally.Our experiments on two mainstream multi-document reading comprehension datasets (Natural Questions and Triv-iaQA) demonstrate that our method can provide consistent improvements over previous state-of-the-art methods.
Liang Wen, Houfeng Wang, Yingwei Luo, Xiaolin Wang 0001
EMNLP3
2022 A Question-Oriented Propagation Network for News Reading Comprehension
abstract
Machine reading comprehension of news articles remains to be a challenging task since the lengths of its context documents are long. Such reading comprehension task usually requires document-level language understanding while state-of-the-art, pretrained question answering models can only encode sequences with a predefined length limit. In this paper, we propose a novel Question-Oriented Propagation Network (QOPN) model for such task. Specifically, our proposed QOPN first uses a context encoding module to find local question-related clues. Then, it employs a multi-step reasoning module to aggregate question-focused information for iterative reasoning. The novel design put emphasis on capturing question-related information and allow long-range information integration, which is especially beneficial for long-context reading comprehension task. Experiments on two challenging machine comprehension datasets show that the proposed QOPN significantly outperforms previous state-of-the-art models.
Liang Wen, Houfeng Wang, Dehong Ma, Yingwei Luo, Xiaolin Wang 0001, Daiting Shi, Zhicong Cheng, Dawei Yin 0001
ICASSP5
2022 Tear Up the Bubble Boom: Lessons Learned From a Deep Learning Research and Development Cluster
abstract
With the proliferation of deep learning, there exists a strong need to efficiently operate GPU clusters for deep learning production in giant AI companies, as well as for research and development (R&D) in small-sized research institutes and universities. Existing works have performed thorough trace analysis on large-scale production-level clusters in giant companies, which discloses the characteristics of deep learning production jobs and motivates the design of scheduling frameworks. However, R&D clusters significantly differ from production-level clusters in both job properties and user behaviors, calling for a different scheduling mechanism. In this paper, we present a detailed workload characterization of an R&D cluster, CloudBrain-I, in a research institute, Peng Cheng Laboratory. After analyzing the fine-grained resource utilization, we discover a severe problem for R&D clusters, resource underutilization, which is especially important in R&D clusters while not characterised by existing works. We further investigate two specific underutilization phenomena and conclude several implications and lessons on R&D cluster scheduling. The traces will be open-sourced to motivate further studies in the community.
Zehua Yang, Zhisheng Ye 0002, Tianhao Fu, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Tianwei Zhang 0004
ICCD6
2022 Exploring GNN based program embedding technologies for binary related tasks
abstract
With the rapid growth of program scale, program analysis, maintenance and optimization become increasingly diverse and complex. Applying learning-assisted methodologies onto program analysis has attracted ever-increasing attention. However, a large number of program factors including syntax structures, semantics, running platforms and compilation configurations block the effective realization of these methods. To overcome these obstacles, existing works prefer to be on a basis of source code or abstract syntax tree, but unfortunately are sub-optimal for binary-oriented analysis tasks closely related to the compilation process. To this end, we propose a new program analysis approach that aims at solving program-level and procedure-level tasks with one model, by taking advantage of the great power of graph neural networks from the level of binary code. By fusing the semantics of control flow graphs, data flow graphs and call graphs into one model, and embedding instructions and values simultaneously, our method can effectively work around emerging compilation-related problems. By testing the proposed method on two tasks, binary similarity detection and dead store prediction, the results show that our method is able to achieve as high accuracy as 83.25%, and 82.77%.
Pengcheng Li 0001, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003
ICPC3
2022 Graph Neural Networks Based Memory Inefficiency Detection Using Selective Sampling
abstract
Production software of data centers oftentimes suffers from unnecessary memory inefficiencies caused by inappropriate use of data structures, conservative compiler optimizations, and so forth. Nevertheless, whole-program monitoring tools often incur incredibly high overhead due to fine-grained memory access instrumentation. Consequently, the fine-grained monitoring tools are not viable for long-running, large-scale data center applications due to strict latency criteria (e.g., service-level agreement or SLA). To this end, this work presents a novel learning-aided system, namely Puffin, to identify three kinds of unnecessary memory operations including dead stores, silent loads and silent stores, by applying gated graph neural networks onto fused static and dynamic program semantics with respect to relative positional embedding. To deploy the system in large-scale data centers, this work explores a sampling-based detection infrastructure with high efficacy and negligible overhead. We evaluate Puffin upon the well-known SPEC CPU 2017 benchmark suite for four compilation options. Experimental results show that the proposed method is able to capture the three kinds of memory inefficiencies with as high accuracy as 96% and a reduced checking overhead by$5.66\times$over the state-of-the-art tool.
Pengcheng Li 0001, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Xu Liu 0001
SC3
2022 Accelerating Address Translation for Virtualization by Leveraging Hardware Mode
abstract
The overhead of memory virtualization remains nontrivial. The traditional shadow paging (TSP) resorts to a shadow page table (SPT) to achieve the native page walk speed, but page table updates require hypervisor interventions. Alternatively, nested paging enables low-overhead page table updates, but utilizes the hardware MMU to perform a long-latency two-dimensional page walk. This paper proposes new memory virtualization solutions based on hardware (machine) mode—the highest CPU privilege level in some architectures like Sunway and RISC-V. A programming interface, running in hardware mode, enables software-implementation of hardware support functions. We first proposeSoftware-based Nested Paging (SNP), which extends the software MMU to perform a two-dimensional page walk in hardware mode. Second, we presentSwift Shadow Paging (SSP), which accomplishes page table synchronization by intercepting TLB flushing in hardware mode. Finally we proposeAccelerated Shadow Paging (ASP)combining SSP and SNP. ASP handles the last-level SPT page faults by walking two-dimensional page tables in hardware mode, which eliminates most hypervisor interventions. This paper systematically compares multiple memory virtualization models by analyzing their designs and evaluating their performance both on a real system and a simulator. The experiments show that the virtualization overhead of ASP is less than 4.5% for all workloads.
Sai Sha, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003
IEEE Trans. Computers3
2022 Astraea: A Fair Deep Learning Scheduler for Multi-Tenant GPU Clusters
abstract
Modern GPU clusters are designed to support distributed Deep Learning jobs from multiple tenants concurrently. Each tenant may have varied and dynamic resource demands. Unfortunately, existing GPU schedulers fail to thoroughly consider the fairness among the tenants and jobs, which can result in unbalanced resource allocation and unfair user experience. In this article, we present an efficient solution to provide strong fairness while maintaining high scheduling effectiveness in multi-tenant GPU clusters. First, we introduce a novel Long-Term GPU-time Fairness metric, which can comprehensively evaluate the fairness at both the tenant and job levels, based on both the temporal and spatial impacts of resource allocation. Second, we design a new and practical GPU scheduler,Astraea, to enforce the desired fairness among tenants and jobs. Large-scale evaluations show thatAstraeacan improve tenant fairness by up to 9.42× compared to state-of-the-art schedulers, without sacrificing the average job completion time.
Zhisheng Ye 0002, Peng Sun 0006, Wei Gao 0064, Tianwei Zhang 0004, Xiaolin Wang 0001, Shengen Yan, Yingwei Luo
IEEE Trans. Parallel Distributed Syst.7
2021 GRAPHSPY: Fused Program Semantic Embedding through Graph Neural Networks for Memory Efficiency
abstract
Production software oftentimes suffers from unnecessary memory inefficiencies caused by inappropriate use of data structures, programming abstractions, or conservative compiler optimizations. Unfortunately, existing works often adopt a whole-program fine-grained monitoring method incurring incredibly high overhead. This work proposes a learning-aided approach to identify unnecessary memory operations, by applying several prevalent graph neural network models to extract program semantics with respect to program structure, execution semantics and dynamic states. Results show that the proposed approach captures memory inefficiencies with high accuracy of 95.27% and only around 17% overhead of the state-of-the-art.
Pengcheng Li 0001, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003
DAC3
2021 An Edge-Fencing Strategy for Optimizing SSSP Computations on Large-Scale Graphs
abstract
The Single-Source Shortest Path (SSSP) problem is to compute the shortest distances in a weighted graph from a source vertex to every other vertex. This paper focuses on parallel efficiency and scalability of SSSP computations on large-scale graphs. We propose an edge-fencing strategy to customize a SSSP algorithm's schedule for every SSSP computation, and devise a path-centric SSSP algorithm with this strategy. This strategy aims at reducing both the relaxed edges and relaxations repeated on each edge. It exploits a few fence values to select the relaxed edges and schedule edge relaxations according to lengths of the created paths. The path-centric algorithm works on a hierarchical graph model, and exploits the edge-fencing strategy to schedule edge relaxations in parallel settings. The hierarchical graph model quantifies the length distribution of shortest paths in large-scale graphs, provides appropriate fence values for every SSSP computation. The algorithm was evaluated on a wide range of synthetic graphs and real-world graphs. The experimental results suggest that our algorithm is efficient and scalable for graphs with skewed degree distributions, and its performance is relatively insensitive to the hierarchical graph model's accuracy.
Huashan Yu, Xiaolin Wang 0001, Yingwei Luo
ICPP3
2021 Swift shadow paging (SSP): no write-protection but following TLB flushing
abstract
Virtualization is a key technique for supporting cloud services and memory virtualization is a major component of virtualization technology. Common memory virtualization mechanisms include shadow paging and hardware-assisted paging. The shadow paging model needs to synchronize shadow/guest page tables whenever there is a guest page table update. In the design of traditional shadow paging (TSP), the guest page table pages are write-protected so the updates can be intercepted by the hypervisor to ensure synchronization. Frequent page table updates cause lots of VM_Exits. Researchers have developed hardware-assisted paging to eliminate this overhead. However, address translation needs to walk a two-dimensional page table. This design significantly increases the overhead of page walk.
Sai Sha, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003
VEE3
2021 Penalty- and Locality-aware Memory Allocation in Redis Using Enhanced AET
abstract
Due to large data volume and low latency requirements of modern web services, the use of an in-memory key-value (KV) cache often becomes an inevitable choice (e.g., Redis and Memcached). The in-memory cache holds hot data, reduces request latency, and alleviates the load on background databases. Inheriting from the traditional hardware cache design, many existing KV cache systems still use recency-based cache replacement algorithms, e.g., least recently used or its approximations. However, the diversity of miss penalty distinguishes a KV cache from a hardware cache. Inadequate consideration of penalty can substantially compromise space utilization and request service time. KV accesses also demonstrate locality, which needs to be coordinated with miss penalty to guide cache management. In this article, we first discuss how to enhance the existing cache model, the Average Eviction Time model, so that it can adapt to modeling a KV cache. After that, we apply the model to Redis and propose pRedis, Penalty- and Locality-aware Memory Allocation in Redis, which synthesizes data locality and miss penalty, in a quantitative manner, to guide memory allocation and replacement in Redis. At the same time, we also explore the diurnal behavior of a KV store and exploit long-term reuse. We replace the original passive eviction mechanism with an automatic dump/load mechanism, to smooth the transition between access peaks and valleys. Our evaluation shows that pRedis effectively reduces the average and tail access latency with minimal time and space overhead. For both real-world and synthetic workloads, our approach delivers an average of 14.0%∼52.3% latency reduction over a state-of-the-art penalty-aware cache management scheme, Hyperbolic Caching (HC), and shows more quantitative predictability of performance. Moreover, we can obtain even lower average latency (1.1%∼5.5%) when dynamically switching policies between pRedis and HC.
Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003
ACM Trans. Storage3
2020 Huge Page Friendly Virtualized Memory Management
Sai Sha, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003
J. Comput. Sci. Technol.3
2019 pRedis: Penalty and Locality Aware Memory Allocation in Redis
abstract
Due to large data volume and low latency requirements of modern web services, the use of in-memory key-value (KV) cache often becomes an inevitable choice (e.g. Redis and Memcached). The in-memory cache holds hot data, reduces request latency, and alleviates the load on background databases. Inheriting from the traditional hardware cache design, many existing KV cache systems still use recency-based cache replacement algorithms, e.g., LRU or its approximations. However, the diversity of miss penalty distinguishes a KV cache from a hardware cache. Inadequate consideration of penalty can substantially compromise space utilization and request service time. KV accesses also demonstrate locality, which needs to be coordinated with miss penalty to guide cache management.
Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003
SoCC2
2019 EMBA: Efficient Memory Bandwidth Allocation to Improve Performance on Intel Commodity Processor
abstract
On multi-core processors, contention on shared resources such as the last level cache (LLC) and memory bandwidth may cause serious performance degradation, which makes efficient resource allocation a critical issue in data centers. Intel recently introduces Memory Bandwidth Allocation (MBA) technology on its Xeon scalable processors, which makes it possible to allocate memory bandwidth in a real system. However, how to make the most of MBA to improve system performance remains an open question. In this work, (1) we formulate a quantitative relationship between a program's performance and its LLC occupancy and memory request rate on commodity processors. (2) Guided by the performance formula, we propose a heuristic bound-aware throttling algorithm to improve system performance and (3) we further develop a hierarchical clustering method to improve the algorithm's efficiency. (4) We implement these algorithms in EMBA, a low-overhead dynamic memory bandwidth scheduling system to improve performance on Intel commodity processors. The results show that, when multiple programs run simultaneously on a multi-core processor whose memory bandwidth is saturated, the programs with high memory bandwidth demand usually use bandwidth inefficiently compared with programs with medium memory bandwidth demand from the perspective of CPU performance. By slightly throttling the former's bandwidth, we can significantly improve the performance of the latter. On average, we improve system performance by 36.9% at the expense of 8.6% bandwidth utilization rate.
Yaocheng Xiang, Chencheng Ye 0001, Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003
ICPP4
2018 DCAPS: dynamic cache allocation with partial sharing
abstract
In a multicore system, effective management of shared last level cache (LLC), such as hardware/software cache partitioning, has attracted significant research attention. Some eminent progress is that Intel introduced Cache Allocation Technology (CAT) to its commodity processors recently. CAT implements way partitioning and provides software interface to control cache allocation. Unfortunately, CAT can only allocate at way level, which does not scale well for a large thread or program count to serve their various performance goals effectively. This paper proposes Dynamic Cache Allocation with Partial Sharing (DCAPS), a framework that dynamically monitors and predicts a multi-programmed workload's cache demand, and reallocates LLC given a performance target. Further, DCAPS explores partial sharing of a cache partition among programs and thus practically achieves cache allocation at a finer granularity. DCAPS consists of three parts: (1) Online Practical Miss Rate Curve (OPMRC), a low-overhead software technique to predict online miss rate curves (MRCs) of individual programs of a workload; (2) a prediction model that estimates the LLC occupancy of each individual program under any CAT allocation scheme; (3) a simulated annealing algorithm that searches for a near-optimal CAT scheme given a specific performance goal. Our experimental results show that DCAPS is able to optimize for a wide range of performance targets and can scale to a large core count.
Yaocheng Xiang, Xiaolin Wang 0001, Zihui Huang, Yingwei Luo, Zhenlin Wang 0003
EuroSys5
2018 Get Out of the Valley: Power-Efficient Address Mapping for GPUs
abstract
GPU memory systems adopt a multi-dimensional hardware structure to provide the bandwidth necessary to support 100s to 1000s of concurrent threads. On the software side, GPU-compute workloads also use multi-dimensional structures to organize the threads. We observe that these structures can combine unfavorably and create significant resource imbalance in the memory subsystem - causing low performance and poor power-efficiency. The key issue is that it is highly application-dependent which memory address bits exhibit high variability. To solve this problem, we first provide an entropy analysis approach tailored for the highly concurrent memory request behavior in GPU-compute workloads. Our window-based entropy metric captures the information content of each address bit of the memory requests that are likely to co-exist in the memory system at runtime. Using this metric, we find that GPU-compute workloads exhibit entropy valleys distributed throughout the lower order address bits. This indicates that efficient GPU-address mapping schemes need to harvest entropy from broad address-bit ranges and concentrate the entropy into the bits used for channel and bank selection in the memory subsystem. This insight leads us to propose the Page Address Entropy (PAE) mapping scheme which concentrates the entropy of the row, channel and bank bits of the input address into the bank and channel bits of the output address. PAE maps straightforwardly to hardware and can be implemented with a tree of XOR-gates. PAE improves performance by 1.31X and power-efficiency by 1.25X compared to state-of-the-art permutation-based address mapping.
Xia Zhao 0004, Magnus Jahre, Zhenlin Wang 0003, Xiaolin Wang 0001, Yingwei Luo, Lieven Eeckhout
ISCA6
2018 Fast Miss Ratio Curve Modeling for Storage Cache
abstract
The reuse distance (least recently used (LRU) stack distance) is an essential metric for performance prediction and optimization of storage cache. Over the past four decades, there have been steady improvements in the algorithmic efficiency of reuse distance measurement. This progress is accelerating in recent years, both in theory and practical implementation. In this article, we present a kinetic model of LRU cache memory, based on the average eviction time (AET) of the cached data. The AET model enables fast measurement and use of low-cost sampling. It can produce the miss ratio curve in linear time with extremely low space costs. On storage trace benchmarks, AET reduces the time and space costs compared to former techniques. Furthermore, AET is a composable model that can characterize shared cache behavior through sampling and modeling individual programs or traces.
Xiameng Hu, Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003, Chen Ding 0001, Chencheng Ye 0001
ACM Trans. Storage4
2017 POSTER: BACM: Barrier-Aware Cache Management for Irregular Memory-Intensive GPGPU Workloads
abstract
General-purpose workloads running on modern graphics processing units (GPGPUs) rely on hardware-based barriers to synchronize warps within a thread block (TB). However, imbalance may exist before reaching a barrier if a GPGPU workload contains irregular memory accesses, i.e., some warps may be critical while others may not. Ideally, cache space should be reserved for the critical warps. Unfortunately, current cache management policies are unaware of the existence of barriers and critical warps, which significantly limits the performance of irregular memory-intensive GPGPU workloads.In this work, we propose Barrier-Aware Cache Management (BACM), which is built on top of two underlying policies: a greedy policy and a friendly policy. The greedy policy does not allow non-critical warps to allocate cache lines in the L1 data cache; only critical warps can. The friendly policy allows non-critical warps to allocate cache lines but only over invalid or lower-priority cache lines. Based on the L1 data cache hit rate of non-critical warps, BACM dynamically chooses between the greedy and friendly policies. By doing so, BACM reserves more cache space to accelerate critical warps, thereby improving overall performance. Experimental results show that BACM achieves an average performance improvement of 24% and 20% compared to the GTO and BAWS policies, respectively. BACM's hardware cost is limited to 96 bytes per streaming multiprocessor.
Xia Zhao 0004, Zhibin Yu 0001, Zhenlin Wang 0003, Xiaolin Wang 0001, Yingwei Luo, Lieven Eeckhout
PACT6
2017 BACM: Barrier-Aware Cache Management for Irregular Memory-Intensive GPGPU Workloads
abstract
General-purpose workloads running on modern graphics processing units rely on hardware-based barriers to synchronize warps within a thread block (TB). However, imbalance may exist before reaching a barrier if a GPGPU workload contains irregular memory accesses, i.e., some warps may be critical while others may not. Ideally, cache space should be reserved for the critical warps. Unfortunately, current cache management policies are unaware of the existence of barriers and critical warps, which significantly limits the performance of irregular memory-intensive GPGPU workloads. In this paper, we propose Barrier-Aware Cache Management (BACM) which is built on top of two underlying policies: a greedy policy and a friendly policy. The greedy policy does not allow non-critical warps to allocate cache lines in the L1 data cache; only critical warps can. The friendly policy allows non-critical warps to allocate cache lines but only over invalid or lower-priority cache lines. BACM dynamically chooses between the greedy and friendly policies based on the L1 data cache hit rate for the non-critical warps. By doing so, BACM reserves more cache space to accelerate critical warps, thereby improving overall performance. Experimental results show that BACM achieves an average performance improvement of 24% and 20% compared to the GTO and BAWS policies, respectively. BACM's hardware cost is limited to 96 bytes per streaming multiprocessor.
Xia Zhao 0004, Zhibin Yu 0001, Zhenlin Wang 0003, Xiaolin Wang 0001, Yingwei Luo, Lieven Eeckhout
ICCD6
2017 Evaluating the impacts of hugepage on virtual machines
Xiaolin Wang 0001, Taowei Luo, Zhenlin Wang 0003, Yingwei Luo
Sci. China Inf. Sci.5
2017 Optimizing Locality-Aware Memory Management of Key-Value Caches
abstract
The in-memory cache system is a performance-critical layer in today's web server architectures. Memcached is one of the most effective, representative, and prevalent among such systems. An important problem is on its memory allocation. The default design does not make the best use of the memory. It is unable to adapt when the demand changes, a problem known as slab calcification. This paper introduces locality-aware memory allocation (LAMA), which addresses the problem by first analyzing the locality of Memcached's requests and then reassigning slabs to minimize the miss ratio or the average response time. By evaluating LAMA using various industry and academic workloads, the paper shows that LAMA outperforms existing techniques in the steady-state performance, the speed of convergence, and the ability to adapt to request pattern changes, and overcome slab calcification. The new solution is close to optimal, achieving over 98 percent of the theoretical potential. Furthermore, LAMA can also be adopted in resource partitioning to guarantee quality-of-service (QoS).
Xiameng Hu, Xiaolin Wang 0001, Yingwei Luo, Chen Ding 0001, Song Jiang 0001, Zhenlin Wang 0003
IEEE Trans. Computers4
2017 Optimal Symbiosis and Fair Scheduling in Shared Cache
abstract
On multi-core processors, applications are run sharing the cache. This paper presents optimization theory to co-locate applications to minimize cache interference and maximize performance. The theory precisely specifies MRC-based composition, optimization, and correctness conditions. The paper also presents a new technique called footprint symbiosis to obtain the best shared cache performance underfair CPU allocation as well as a new sampling technique which reduces the cost of locality analysis. When sampling and optimization are combined, the paper shows that it takes less than 0.1 second analysis per program to obtain a co-run that is within 1.5 percent of the best possible performance. In an exhaustive evaluation with 12,870 tests, the best prior work improves co-run performance by 56 percent on average. The new optimization improves it by another 29 percent. Without single co-run test, footprint symbiosis is able to choose co-run choices that are just 8 percent slower than the best co-run solutions found with exhaustive testing.
Xiameng Hu, Xiaolin Wang 0001, Yechen Li, Yingwei Luo, Chen Ding 0001, Zhenlin Wang 0003
IEEE Trans. Parallel Distributed Syst.4
2016 Barrier-Aware Warp Scheduling for Throughput Processors
abstract
Parallel GPGPU applications rely on barrier synchronization to align thread block activity. Few prior work has studied and characterized barrier synchronization within a thread block and its impact on performance. In this paper, we find that barriers cause substantial stall cycles in barrier-intensive GPGPU applications although GPGPUs employ lightweight hardware-support barriers. To help investigate the reasons, we define the execution between two adjacent barriers of a thread block as a warp-phase. We find that the execution progress within a warp-phase varies dramatically across warps, which we call warp-phase-divergence. While warp-phase-divergence may result from execution time disparity among warps due to differences in application code or input, and/or shared resource contention, we also pinpoint that warp-phase-divergence may result from warp scheduling.
Zhibin Yu 0001, Lieven Eeckhout, Vijay Janapa Reddi, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Cheng-Zhong Xu 0001
ICS5
2016 Kinetic Modeling of Data Eviction in Cache
Xiameng Hu, Xiaolin Wang 0001, Yingwei Luo, Chen Ding 0001, Zhenlin Wang 0003
USENIX ATC4
2016 A survey of cloud resource management for complex engineering applications
Haibao Chen, Song Wu 0001, Hai Jin 0001, Jidong Zhai, Yingwei Luo, Xiaolin Wang 0001
Frontiers Comput. Sci.6
2016 Dynamic Memory Balancing for Virtualization
abstract
Allocating memory dynamically for virtual machines (VMs) according to their demands provides significant benefits as well as great challenges. Efficient memory resource management requires knowledge of the memory demands of applications or systems at runtime. A widely proposed approach is to construct a miss ratio curve (MRC) for a VM, which not only summarizes the current working set size (WSS) of the VM but also models the relationship between its performance and the target memory allocation size. Unfortunately, the cost of monitoring and maintaining the MRC structures is nontrivial. This article first introduces a low-cost WSS tracking system with effective optimizations on data structures, as well as an efficient mechanism to decrease the frequency of monitoring. We also propose a Memory Balancer (MEB), which dynamically reallocates guest memory based on the predicted WSS. Our experimental results show that our prediction schemes yield a high accuracy of 95.2% and low overhead of 2%. Furthermore, the overall system throughput can be significantly improved with MEB, which brings a speedup up to 7.4 for two to four VMs and 4.54 for an overcommitted system with 16 VMs.
Xiaolin Wang 0001, Fang Hou 0005, Yingwei Luo, Zhenlin Wang 0003
ACM Trans. Archit. Code Optim.4
2015 Improving TLB Performance by Increasing Hugepage Ratio
abstract
Linux supports transparent huge page since 2.6.38. It can automatically map huge pages. But this implementation fails to adjust to page alignment in memory allocation and thus cannot use huge page in some situations. The design is not efficient. Our work aims to increase huge page allocation, so as to improve the utilization ratio of huge page and overall performance. The experimental results show that the optimization almost reaches the upper bound of huge page utilization. This software approach delivers a notable performance improvement for a few benchmarks with moderate overhead in physical memory consumption.
Taowei Luo, Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003
CCGRID4
2015 Optimal Footprint Symbiosis in Shared Cache
abstract
On multicore processors, applications are run sharing the cache. This paper presents online optimization to collocate applications to minimize cache interference to maximize performance. The paper formulates the optimization problem and solution, presents a new sampling technique for locality analysis and evaluates it in an exhaustive test of 12,870 cases. For locality analysis, previous sampling was two orders of magnitude faster than full-trace analysis. The new sampling reduces the cost by another two orders of magnitude. The best prior work improves co-run performance by 56% on average. The new optimization improves it by another 29%. When sampling and optimization are combined, the paper shows that it takes less than 0.1 second analysis per program to obtain a co-run that is within 1.5% of the best possible performance.
Xiaolin Wang 0001, Yechen Li, Yingwei Luo, Xiameng Hu, Jacob Brock, Chen Ding 0001, Zhenlin Wang 0003
CCGRID3
2015 Optimal Cache Partition-Sharing
abstract
When a cache is shared by multiple cores, its space may be allocated either by sharing, partitioning, or both. We call the last case partition-sharing. This paper studies partition-sharing as a general solution, and presents a theory an technique for optimizing partition-sharing. We present a theory and a technique to optimize partition sharing. The theory shows that the problem of partition-sharing is reducible to the problem of partitioning. The technique uses dynamic programming to optimize partitioning for overall miss ratio, and for two different kinds of fairness. Finally, the paper evaluates the effect of optimal cache sharing and compares it with conventional solutions for thousands of 4-program co-run groups, with nearly 180 million different ways to share the cache by each co-run group. Optimal partition-sharing is on average 26% better than free-for-all sharing, and 98% better than equal partitioning. We also demonstrate the trade-off between optimal partitioning and fair partitioning.
Jacob Brock, Chencheng Ye 0001, Chen Ding 0001, Yechen Li, Xiaolin Wang 0001, Yingwei Luo
ICPP6
2015 LAMA: Optimized Locality-aware Memory Allocation for Key-value Cache
Xiameng Hu, Xiaolin Wang 0001, Yechen Li, Yingwei Luo, Chen Ding 0001, Song Jiang 0001, Zhenlin Wang 0003
USENIX ATC5
2014 Performance Metrics and Models for Shared Cache
Chen Ding 0001, Xiaoya Xiang, Bin Bao, Hao Luo 0007, Yingwei Luo, Xiaolin Wang 0001
J. Comput. Sci. Technol.5
2013 Towards Eliminating Memory Virtualization Overhead
Xiaolin Wang 0001, Lingmei Weng, Zhenlin Wang 0003, Yingwei Luo
APPT4
2013 Who decides migration? A migration lock mechanism for virtual machines
abstract
Migration of virtual machines is an important feature for the management of a virtualized environment. Current strategies for managing migration consider more of resource scheduling and system maintenance, ignoring possible constraints on migration from virtual machines and applications. This paper introduces a migration locking mechanism and its implementation on Xen. The virtual machine migration locking mechanism provides a standard interface for virtual machine users to control migration status of their virtual machines according to their own wishes, such as security, hardware or performance demands and so on. But forced migration locking by the end users or applications could conflict with the management decisions by managers or the virtual machine monitors. How to coordinate different requirements between virtual machine users and managers needs further discussion.
Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003
CNSM2
2013 Failure Recovery: When the Cure Is Worse Than the Disease
Sean McDirmid, Mao Yang 0004, Li Zhuang, Yingwei Luo, Tom Bergan, Madan Musuvathi, Zheng Zhang 0001, Lidong Zhou
HotOS6
2013 Revisiting memory management on virtualized environments
abstract
With the evolvement of hardware, 64-bit Central Processing Units (CPUs) and 64-bit Operating Systems (OSs) have dominated the market. This article investigates the performance of virtual memory management of Virtual Machines (VMs) with a large virtual address space in 64-bit OSs, which imposes different pressure on memory virtualization than 32-bit systems. Each of the two conventional memory virtualization approaches, Shadowing Paging (SP) and Hardware-Assisted Paging (HAP), causes different overhead for different applications. Our experiments show that 64-bit applications prefer to run in a VM using SP, while 32-bit applications do not have a uniform preference between SP and HAP. In this article, we trace this inconsistency between 32-bit applications and 64-bit applications to its root cause through a systematic empirical study in Linux systems and discover that the major overhead of SP results from memory management in the 32-bit GNU C library ( glibc ). We propose enhancements to the existing memory management algorithms, which substantially reduce the overhead of SP. Based on the evaluations using SPEC CPU2006, Parsec 2.1, and cloud benchmarks, our results show that SP, with the improved memory allocators, can compete with HAP in almost all cases, in both 64-bit and 32-bit systems. We conclude that without a significant breakthrough in HAP, researchers should pay more attention to SP, which is more flexible and cost effective.
Xiaolin Wang 0001, Lingmei Weng, Zhenlin Wang 0003, Yingwei Luo
ACM Trans. Archit. Code Optim.4
2012 Live Migrating the Virtual Machine Directly Accessing a Physical NIC
abstract
This paper proposes a dynamic physical-virtual NIC switching. A virtual machine directly accessing the physical NIC can switch to use the virtual NIC if it is to be migrated. This solution is implemented in network configuration level in Guest OS drivers, so it can be used in different VMMs and with different NICs. The experiments illustrate that our solution does not interrupt the network connections from clients' perspective. After the virtual machine switches to use the virtual NIC, it can be migrated and the switching will not affect the migration performance, including the downtime during the migration.
Xiaolin Wang 0001, Yingwei Luo, Xiaoming Li 0001, Zhenlin Wang 0003
APSCC3
2012 A model contract and model integration language for integrating geography models in distributed environment
abstract
This paper presents the model contract and the model integration language, which is used for sharing and reusing geography models in a distributed geography modeling environment. The model contract mainly consists of three parts, the model external description information, the model internal structural information and the model execution flow information. Geographer can structure the model contract to invoke shared models and method to simulate the geography phenomenon using the model integration language. The model integration language is a kind of markup language to express the model contract.
Xiaolin Wang 0001, Yingwei Luo
IGARSS3
2012 Design model execution engine based on web services for distributed geography modeling environment
abstract
The design of execution engine based on web services for distributed geography modeling environment relies on model contract, model management environment, execution node, data center and control node. The execution procedure of geography model execution engine would be like this: first, geographers construct model contract according to the rule of model contract, and then submit the model contract into control node; Second, the control node parsers the model contract, queries the geography model deployment information, entry point and model input/output information from model management environment, keeps them in memory; Third, control node prepares model data and calls the models which distributed on each execution node for simulation; Finally, the environment returns the result to geographers. The model execution engine adopts multi-thread technology, and could handle multi-job submitted by geographers.
Xiaolin Wang 0001, Hongqiang Mao, Yingwei Luo
IGARSS3
2012 A Dynamic Cache Partitioning Mechanism under Virtualization Environment
abstract
Cache sharing among multiple computing units on chip is common in today's multi-core processors, and a lot of research has focused on the effective management of shared cache. A software management method called page coloring is commonly used to divide the cache among different applications competing for the same cache entries. Both static and dynamic cache partition mechanism have been implemented in operating system or user-space level. However, few efforts have been made under virtualization environments. Our previous work has provided a static cache partition method based on page coloring in Xen, following that, a dynamic cache partition mechanism called Colored Page Migration (CoPaM) is presented in this paper.
Xiaolin Wang 0001, Yechen Li, Yingwei Luo, Xiaoming Li 0001, Zhenlin Wang 0003
TrustCom4
2012 Dynamic cache partitioning based on hot page migration
Xiaolin Wang 0001, Yechen Li, Zhenlin Wang 0003, Yingwei Luo, Xiaoming Li 0001
Frontiers Comput. Sci.5
2011 Model semantic network for massive spatial information
abstract
Spatial information is now increasing continuously and is available via the Internet. But facing the abundant spatial information, people recognize many agonizing problems, such as how can spatial information cooperate with each other to solve users' task and how do users know where spatial information is located, what kinds of spatial information can be used, and the way of how to use them. The only way to remedy those problems is to use spatial metadata. Spatial metadata is the description of spatial information and its associated information, which give a semantic expatiation for spatial information. Spatial metadata is now developed as an indispensable and powerful tool for data finding, data exchanging, data managing and data utilization. Currently, there are some standards for spatial metadata such as FGDC and ISO/TC211. But the existing spatial metadata is largely for people to use, and only describes spatial information itself. In this paper, spatial metadata is extended to describe the relations among different spatial information and the distribution of spatial information on the Internet. Based on the extensive spatial metadata, a semantic network model for massive spatial information is presented to support the collaboration among spatial information and make the navigation of spatial information more fast and exact. At last, XML technology is adopted to represent the semantic network for spatial information.
Xiaolin Wang 0001, Yingwei Luo
IGARSS2
2011 Managing and integrating geography models in distributed environment
abstract
For the current "Model islands" problems which occur in the process of the geography modeling and model sharing, we propose an idea of geography model sharing and reusing method which based on the metadata standard in distributed environment. This article mainly solves how to describe and manage the geography model to achieve the model sharing and reuse. We work on the design of the metadata standard, the integrating standard and the managing environment of the geography model. It is expected to provide a convenient platform for the geographers, so that they can easily reuse any model, which means the heterogeneous models can be well shared.
Xiaolin Wang 0001, Yingwei Luo
IGARSS4
2011 Sharing and reusing geography models via model execution engine
abstract
This article analyzes the sharing and reuse of the geography model in the distributed environment based on the methods of geography model integration and framework. We use the model contract to describe the integration of geography model, and design a model contract language to express. We also design the model execution engine architecture which can execute the model contract language. This is the core function to achieve the sharing and reuse of the geography model in distributed environment.
Xiaolin Wang 0001, Hongqiang Mao, Yingwei Luo
IGARSS5
2011 Low Cost Working Set Size Tracking
Weiming Zhao, Xinxin Jin, Zhenlin Wang 0003, Xiaolin Wang 0001, Yingwei Luo, Xiaoming Li 0001
USENIX ATC5
2011 Selective hardware/software memory virtualization
abstract
As virtualization becomes a key technique for supporting cloud computing, much effort has been made to reduce virtualization overhead, so a virtualized system can match its native performance. One major overhead is due to memory or page table virtualization. Conventional virtual machines rely on a shadow mechanism to manage page tables, where a shadow page table maintained by the VMM (Virtual Machine Monitor) maps virtual addresses to machine addresses while a guest maintains its own virtual to physical page table. This shadow mechanism will result in expensive VM exits whenever there is a page fault that requires synchronization between the two page tables. To avoid this cost, both Intel and AMD provide hardware assists, EPT (extended page table) and NPT (nested page table), to facilitate address translation. With the hardware assists, the MMU (Memory Management Unit) maintains an ordinary guest page table that translates virtual addresses to guest physical addresses. In addition, the extended page table as provided by EPT translates from guest physical addresses to host physical or machine addresses. NPT works in a similar style. With EPT or NPT, a guest page fault can be handled by the guest itself without triggering VM exits. However, the hardware assists do have their disadvantage compared to the conventional shadow mechanism -- the page walk yields more memory accesses and thus longer latency. Our experimental results show that neither hardware-assisted paging (HAP) nor shadow paging (SP) can be a definite winner. Despite the fact that in over half of the cases, there is no noticeable gap between the two mechanisms, an up to 34% performance gap exists for a few benchmarks. We propose a dynamic switching mechanism that monitors TLB misses and guest page faults on the fly, and dynam-ically switches between the two paging modes. Our experiments show that this new mechanism can match and, sometimes, even beat the better performance of HAP and SP.
Xiaolin Wang 0001, Jiarui Zang, Zhenlin Wang 0003, Yingwei Luo, Xiaoming Li 0001
VEE4
2011 A Rule-Based Pretreatment Mechanism for Online Mobile Map Data
abstract
In online map service for mobile users, it's necessary to provide different map data according to different application scenarios. For a given scenario, we can preprocess the map data to satisfy the online and mobile requirements. This paper proposes a rule-based pretreatment mechanism for online mobile map data which using a novel mobile map data format Byte-Map. Through the use of pretreatment rules definition and rule engine, we can externalize the decision logic from the business logic in data pretreatment mechanism. We can also separate the data preparer and data pretreatment programmer from each other. In one hand, the data preparer can make definition and configure rules to change the pretreatment procedure, without knowing the program code. In another hand, the data pretreatment programmer can avoid modifying and re-compiling program code for handling map data from different source and type. The rule-based pretreatment mechanism made the online mobile map data pretreatment procedure flexible and extensional.
Xiaolin Wang 0001, Yingwei Luo
VTC Fall3
2010 The design and implementation of GIS applications based on SOA
abstract
With the rapid development of network technologies, SOA (Service Oriented Architecture), which is a methodology of constructing the enterprise distributed software systems, has been widely used nowadays. In this paper we discussed mainly about how to construct a traditional GIS application with the related technologies of SOA. First, we explained the two key problems about module reuse in GIS: encapsulation and combination, and the issues which haven't been addressed till now. Next, we described some simple situations of applying SOA in GIS. At last, we showed PKUMAP, which is a light weighted WebGIS application, using the kernel technologies of SOA, such as SCA, BPEL and so on.
Xiaolin Wang 0001, Xiao Pang, Yingwei Luo
IGARSS3
2010 Web Service encapsulation of fortran-based geographical model
abstract
This paper take the hydrological model SWAT as an example, explores the encapsulation technology which encapsulate a Fortran source code into a Web Service. In the geographical models, there are a great many of global variable parameters and file reading parameters, according to this situation, we put out a method which makes use of XML file to transform data and file pointer parameters. The encapsulation method is the basis of constructing the model library in the distributed geographical model environment.
Xiaolin Wang 0001, Yingwei Luo
IGARSS4
2010 Evaluating and Optimizing I/O Virtualization in Kernel-based Virtual Machine (KVM)
Xiaolin Wang 0001, Rongfeng Lai, Zhenlin Wang 0003, Yingwei Luo, Xiaoming Li 0001
NPC6
2010 LBS-p: A LBS Platform Supporting Online Map Services
abstract
This paper presents LBS-p, a LBS supporting platform, which provides online map service. LBS-p consists of LBS-p Mobile and LBS-p Server. LBS-p Mobile is a Java ME application running on the mobile terminal, which dedicates to the request, management and display of mobile map data. LBS-p Server consists of data pre-processing mechanism, data providing module and LBS-oriented GIS service module. Performance evaluations and a LBS application example show that LBS-p is effective.
Xiaolin Wang 0001, Xiao Pang, Yingwei Luo
VTC Fall3
2010 Byte-Map: A Novel Mobile Map Format Using Two-Byte Coordinates
abstract
This paper presents Byte-Map, which is a novel mobile map format for mobile map service in mobile devices. Byte-Map is a kind of vector format with different blocks through different levels, and map data of Byte-Map is encapsulated in binary stream. The basic cell of Byte-Map is a block, which is fixed in size of 255 units*255 units according to different coordinates systems and thus the coordinates of all the features in a certain block can be encoded with only two bytes. The experiment shows that Byte-Map has good properties in data volume, data decompressing, data parsing, map displaying and memory cost.
Xiaolin Wang 0001, Xiao Pang, Yingwei Luo
VTC Fall3
2010 DMM: A dynamic memory mapping model for virtual machines
Haogang Chen 0002, Xiaolin Wang 0001, Zhenlin Wang 0003, Yingwei Luo, Xiaoming Li 0001
Sci. China Inf. Sci.5
2010 Dynamic memory paravirtualization transparent to guest OS
Xiaolin Wang 0001, Yifeng Sun, Yingwei Luo, Zhenlin Wang 0003, Yu Li 0006, Haogang Chen 0002, Xiaoming Li 0001
Sci. China Inf. Sci.3
2009 Fast Live Cloning of Virtual Machine Based on Xen
abstract
Virtual Machine (VM) cloning is to create a replica of a source virtual machine (parent virtual machine); the replica, also called child virtual machine, owns exactly the same executing status as parent virtual machine. Fast live cloning guarantees that, during the period of cloning, the services running on the parent virtual machine observe no performance degradation. There are three important goals for fast live cloning: reducing the total cloning time, minimizing the suspension time of the parent virtual machine, and maximizing resource sharing between the parent virtual machine and the child virtual machine. This paper exploits Copy-on-Write (CoW) mechanism fully to minimize the time of duplicating active main memory and secondary storage of the parent virtual machine. As a result, out experiments show that the total cloning time of a virtual machine on Xen virtual machine monitor can be confined within several hundred milliseconds and the downtime of parent virtual machine is limited to tens of milliseconds, close to VMM’s scheduling interval. Experiments also show that the parent virtual machine highly shares memory with the child virtual machine as long as both virtual machines continue executing the same applications.
Yifeng Sun, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Haogang Chen 0002, Xiaoming Li 0001
HPCC2
2009 REMOCA: Hypervisor Remote Disk Cache
abstract
In virtual machine (VM) systems, with the increase in the number of VMs and the demands of applications, the main memory is becoming a bottleneck of application performance. To improve paging performance for memory-intensive or I/O-intensive workloads, we propose the hypervisor REMOte disk CAche (REMOCA), which allows a virtual machine to use the memory resources on other physical machines as its cache between its virtual memory and virtual disk devices. The goal of REMOCA is to reduce disk accesses, which is much slower than transferring memory pages over modern interconnect networks. As a result, the average disk I/O latency can be improved. REMOCA is implemented within the hypervisor, by intercepting guest events such as page evictions and disk accesses. This design is transparent to the applications, and is compatible with existing techniques like ballooning and ghost buffer. Moreover, a combination of them can provide a more flexible resource management policy. Our experimental results show that REMOCA can efficiently alleviate the impact of thrashing behavior, and also significantly improve the performance for real-world I/O intensive applications.
Haogang Chen 0002, Xiaolin Wang 0001, Zhenlin Wang 0003, Xinxin Jin, Yingwei Luo, Xiaoming Li 0001
ISPA6
2009 A Simple Cache Partitioning Approach in a Virtualized Environment
abstract
Virtualization is often used in systems for the purpose of offering isolation among applications running in separate virtual machines (VM). Current virtual machine monitors (VMMs) have done a decent job in resource isolation in memory, CPU and I/O devices. However, when looking further into the usage of lower-level shared cache, we notice that one virtual machine’s cache behavior may interfere with another’s due to the uncontrolled cache sharing. In this situation, performance isolation cannot be guaranteed. This paper presents a cache partitioning approach which can be implemented in the VMM. We have implemented this mechanism in Xen VMM using the page coloring technique traditionally applied to the OS. Our VMM-based implementation is fully transparent to the guest OSes. It thus shows the advantages of simplicity and flexibility. Our evaluation shows that our cache partitioning method can work efficiently and improve the performance of co-scheduled applications running within different VMs. In the concurrent workloads selected from the SPEC CPU 2006 benchmarks, our technique achieves a performance improvement by up to 19% for the most sensitive workloads
Xinxin Jin, Haogang Chen 0002, Xiaolin Wang 0001, Zhenlin Wang 0003, Yingwei Luo, Xiaoming Li 0001
ISPA6
2009 A Refined Mobile Map Format and Its Application
Yingwei Luo, Xiaolin Wang 0001, Xiao Pang
SSTD1
2008 A Rule-Based Event Handling Model
abstract
A rule-based event handling model is proposed in the paper, which has a hierarchical architecture with four layers: resource layer, knowledge layer, business layer and representation layer. This model is aimed at distributed information integrating and service scheduling for event handling through rule in knowledge layer. Detail work on knowledge layer is explored, which includes definition of formal business rule, reference to resources in a rule, as well as design and implementation of the rule system. A unified urgent event call-in disposition of city joint emergency response systems is taking as an example of rule-based information integrating and service scheduling.
Yingwei Luo, Xiaolin Wang 0001, Xinpeng Liu 0001, Zhou Xing, Xiao Pang
APSCC1
2008 Live and incremental whole-system migration of virtual machines using block-bitmap
abstract
In this paper, we describe a whole-system live migration scheme, which transfers the whole system run-time state, including CPU state, memory data, and local disk storage, of the virtual machine (VM). To minimize the downtime caused by migrating large disk storage data and keep data integrity and consistency, we propose a three-phase migration (TPM) algorithm. To facilitate the migration back to initial source machine, we use an incremental migration (IM) algorithm to reduce the amount of the data to be migrated. Block-bitmap is used to track all the write accesses to the local disk storage during the migration. Synchronization of the local disk storage in the migration is performed according to the block-bitmap. Experiments show that our algorithms work well even when I/O-intensive workloads are running in the migrated VM. The downtime of the migration is around 100 milliseconds, close to shared-storage migration. Total migration time is greatly reduced using IM. The block-bitmap based synchronization mechanism is simple and effective. Performance overhead of recording all the writes on migrated VM is very low.
Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Yifeng Sun, Haogang Chen 0002
CLUSTER1
2008 ChinaV: Building Virtualized Computing System
abstract
Virtualization technology has attracted much attention in recent years. This paper describes the vision and mission of ChinaV, which is the national fundamental research program for virtualization technology in China. Furthermore, related topics about single host virtualization, multiple VM management schemes and desktop virtualization will be introduced. We first describe a remote memory virtualization scheme and a VCPU management scheme for efficient use of physical resource. Then we describe a novel live VM migration approach based on deterministic replay with execution trace. Multiple VM management schemes are also introduced for multi-VM virtualization. In desktop virtualization field, we present the LVD, a system that combines the virtualization technology and inexpensive personal computers to realize a lightweight virtual desktop system. All of those schemes and systems are good practices of virtualization solution and they have become a strong foundation of our future work.
Hai Jin 0001, Xiaofei Liao, Song Wu 0001, Zhiyuan Shao, Yingwei Luo
HPCC5
2005 Spatial Data Channel in a Mobile Navigation System
Yingwei Luo, Guomin Xiong, Xiaolin Wang 0001, Zhuoqun Xu
ICCSA (2)1
2005 Ontological Model of Event for Integration of Inter-organization Applications
Wenjun Wang 0002, Yingwei Luo, Xinpeng Liu 0001, Xiaolin Wang 0001, Zhuoqun Xu
ICCSA (1)2
2005 XML Approach to Communication Design of WebGIS
Yingwei Luo, Xinpeng Liu 0001, Xiaolin Wang 0001, Zhuoqun Xu
ICWE1
2005 The Study and Application of Crime Emergency Ontology Event Model
Wenjun Wang 0002, Yingwei Luo, Xiaolin Wang 0001, Zhuoqun Xu
KES (4)3
2004 SOM: A Novel Model for Defining Topological Line-Region Relations
Xiaolin Wang 0001, Yingwei Luo, Zhuoqun Xu
ICCSA (3)2
2004 A Component-Based WebGIS Geo-Union
Yingwei Luo, Xiaolin Wang 0001, Zhuoqun Xu
ICWE1
2004 GML Based Ubiquitous WebGIS
Yingwei Luo, Baoqi Huang, Jiangong Xu, Xiaolin Wang 0001, Zhuoqun Xu
SNPD1
2004 Agent-based Spatial Information Collaboration and Parallel Mechanisms
Yingwei Luo, Xiaolin Wang 0001, Zhuoqun Xu
SNPD1
2004 Component-Based WebGIS and Its Spatial Cache Framework
Yingwei Luo, Xiaolin Wang 0001, Zhuoqun Xu
WAIM1
2003 Extension of spatial metadata and agent-based spatial Data navigation mechanism
abstract
Fast navigation to distributed spatial data has been the keystone of distributed GIS. In this paper, a hierarchical spatial metadata Database framework is presented based on the existing spatial metadata standard. This framework can efficiently organize the distributed spatial data in network. With the support of the hierarchical spatial metadata Databases, a user-oriented descriptive specification for spatial data requirement is proposed. A map (map-layer) is described as a tuple consisting of four basic elements of . An agent-based searching scheme for spatial query is also introduced, which can satisfy the requirement of fast navigation in terms of spatial query that can often be non-deterministic or with uncertainty.
Yingwei Luo, Xiaolin Wang 0001, Zhuoqun Xu
GIS1
2003 Extension of spatial metadata for navigating distributed spatial data
abstract
Navigating distributed spatial data have been the keystone of distributed GIS. Using spatial metadata is an effective method to locate and access spatial data. In this paper, based on the existing spatial metadata standards, a hierarchical spatial metadata database framework is presented. The framework is extended to describe the distribution of spatial databases, the relations among different spatial databases in network and how to support the fast navigation to spatial data. The framework consists of Global spatial metadata database (Global SMDB), District spatial metadata database (District SMDB) and General spatial metadata database (General SMDB), which can efficiently organize and manage distributed spatial data in network. Spatial data can be navigated quickly with support of the framework.
Yingwei Luo, Xiaolin Wang 0001, Zhuoqun Xu
IGARSS1