VLDB 2026 Research / reviewers in the wild / expert
Xiaolin Wang 0001
dblp:29/4241-1
· DBLP profile ↗
87ranked-venue papers
20as first author
29since 2021 · last 2026
0000-0002-6951-1613ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 46 · 4 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 11 first-authorSoftware engineering, systems software and programming languages · 8 · 3 since 2021Artificial intelligence and machine learning · 6 · 2 since 2021Databases, data management, data science and information retrieval · 5Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlowGPU: Transparent and Efficient GPU Checkpointing and Restore
Zehua Yang, Yonghao Zou, Junyang Zhang 0003, Zhisheng Ye 0002, Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003, Diyu Zhou |
Euro-Par (2) | 7 |
| 2026 | Latency-SLO-Aware Memory Offloading for Large Language Model InferenceabstractOffloading large language models (LLMs) states to host memory during inference promises to reduce operational costs by supporting larger models, longer prompts, and larger batch sizes. However, the design of existing memory offloading mechanisms does not take latency service-level objectives (SLOs) into consideration. As a result, they either lead to frequent SLO violations or underutilize host memory, thereby incurring economic loss and thus defeating the purpose of memory offloading. Chenxiang Ma, Zhisheng Ye 0002, Zehua Yang, Tianhao Fu, Jiaxun Han, Jie Zhang 0048, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Yong Li 0045, Diyu Zhou |
ICS | 9 |
| 2026 | A Comprehensive Study on Solving Memory Bloat Under VirtualizationabstractHuge pages are effective in reducing address translation overhead under virtualization. However, huge pages can lead to the memory bloat problem, which manifests in two primary forms: hot bloat and usage bloat . Hot bloat occurs when accesses to a huge page are heavily skewed towards a small subset of base pages, leading the hypervisor to (mistakenly) classify the entire huge page as hot. Hot bloat undermines several critical virtualization techniques, including tiered memory and page sharing. Usage bloatrefers to the base pages within a huge page that has not yet been allocated, causing virtual machines (VMs) to demand excessive memory. Prior work addressing memory bloat either requires hardware modification or targets a specific scenario and is not applicable to a hypervisor. This article presents HugeScope , a lightweight, effective and generic system that addresses the memory bloat problem under virtualization based on commodity hardware. HugeScope includes an efficient and precise page tracking mechanism, leveraging the other level of indirect memory translation in the hypervisor. HugeScope provides a generic framework to support page splitting and coalescing policies, considering the memory pressure, as well as the recency, frequency, and skewness of page access. Moreover, HugeScope is general and modular. It can not only be easily applied to various scenarios concerning hot bloat , including tiered memory management ( HS-TMM ) and page sharing ( HS-Share ), but also seamlessly expose its capabilities to VMs to address the usage bloat problem ( HS-HP ). Evaluation shows that HugeScope incurs less than 4% overhead, by addressing hot bloat , HS-TMM improves performance by up to 61% over vTMM while HS-Share saves 41% more memory than Ingens while offering comparable performance, and By addressing usage bloat , HS-HP can eliminate excessive memory usage, and achieve performance improvements of up to 11% over HawkEye. Chuandong Li 0004, Dong Liu 0042, Zhihong Xue, Xiaolin Wang 0001, Zhenlin Wang 0003, Yingwei Luo, Diyu Zhou |
ACM Trans. Comput. Syst. | 5 |
| 2026 | One Sketch is Enough: Accurate Per-Flow Tail Latency Estimation With SketchPolymer
Jiarui Guo, Yuqi Dong, Yuhan Wu 0001, Yisen Hong, Xiaolin Wang 0001, Yong Cui 0001, Bin Cui 0001, Tong Yang 0003 |
IEEE Trans. Netw. | 6 |
| 2025 | SPDK+: Low Latency or High Power Efficiency? We Take BothabstractSPDK, as one of the most efficient I/O storage software, is capable of delivering the lowest I/O latency. Unfortunately, the polling mechanism in SPDK wastes tremendous CPU clock cycles, especially under small I/O operations and low queue depths. Although SPDK supports the conventional interrupt method, it does not improve power efficiency under such circumstances. To address this issue, we propose SPDK+, which enables the user interrupt feature in the SPDK to achieve both low latency and high power efficiency. Specifically, SPDK+ employs user interrupt handling to directly process MSI-X interrupts from SSD devices and utilizes user wait instructions during IO wait periods to conserve power. The comprehensive evaluation results show that SPDK+ achieves up to 49.5% power efficiency improvement while keeping the I/O latency almost unchanged compared with SPDK. Endian Li, Shushu Yi, Qiao Li 0001, Diyu Zhou, Zhenlin Wang 0003, Xiaolin Wang 0001, Bo Mao 0003, Yingwei Luo, Ke Zhou 0001, Jie Zhang 0048 |
HotStorage | 7 |
| 2025 | InstAttention: In-Storage Attention Offloading for Cost-Effective Long-Context LLM InferenceabstractThe widespread of Large Language Models (LLMs) marks a significant milestone in generative AI. Nevertheless, the increasing context length and batch size in offline LLM inference escalate the memory requirement of the key-value (KV) cache, which imposes a huge burden on the GPU VRAM, especially for resource-constrained scenarios (e.g., edge computing). Several cost-effective solutions leverage host memory or SSDs to reduce storage costs for offline inference scenarios and improve the throughput. Nevertheless, they suffer from significant performance penalties imposed by intensive KV cache accesses due to limited PCIe bandwidth. To address these issues, we propose InstAttention, a novel LLM inference system that offloads the most performance-critical computation (i.e., attention in decoding phase) and data (i.e., KV cache) parts to Computational Storage Drives (CSDs), which minimize the enormous KV transfer overheads. InstAttention designs a dedicated flashaware in-storage attention engine with KV cache management mechanisms to exploit the high internal bandwidths of CSDs instead of being limited by the PCIe bandwidth. The optimized P2P transmission between GPU and CSDs further reduces data migration overheads. Experimental results demonstrate that for a 13B model using an NVIDIA A6000 GPU, InstAttention improves throughput for long-sequence inference by up to $11.1 \times$, compared to existing SSD-based solutions such as FlexGen. Xiurui Pan, Endian Li, Qiao Li 0001, Shengwen Liang, Yizhou Shan, Ke Zhou 0001, Yingwei Luo, Xiaolin Wang 0001, Jie Zhang 0048 |
HPCA | 8 |
| 2025 | Criticality-Aware Instruction-Centric Bandwidth Partitioning for Data Center ApplicationsabstractTo reduce operational costs, modern data centers co-locate high-priority latency-critical (LC) tasks and low-priority best-effort (BE) tasks on the same physical node to increase resource utilization. However, such co-location leads to contention for memory bandwidth, resulting in priority inversion, where BE tasks severely slow down LC tasks. This priority inversion often leads to violations of the quality of service (QoS) requirements for LC tasks, defeating the purpose of co-location. Prior approaches to this issue either fail to enforce the QoS requirements for LC tasks or underutilize memory bandwidth.We present Pivot, a novel bandwidth partitioning system that overcomes the limitations of prior approaches based on two key insights. First, memory accesses from LC tasks must be prioritized across all the components on the memory path rather than a single component, as done in prior work. Second, only the scheduling of a selective portion of performance-critical loads (i.e., those causing a long stall on the re-order buffer), instead of all memory accesses from LC tasks, should be prioritized. To leverage these insights, Pivot overcomes the key challenge of accurately identifying performance-critical loads while incurring minimal runtime overhead by proposing a two-phase profiling technique. Our extensive evaluation shows that Pivot improves effective machine utilization by up to $\mathbf{3 4. 5 \%}$ while increasing the throughput of the BE applications by up to $2.76 \times$ compared to state-of-the-art approaches. Liren Zhu, Liujia Li, Jie Zhang 0048, Zhenlin Wang 0003, Xiaolin Wang 0001, Yingwei Luo, Diyu Zhou |
HPCA | 8 |
| 2025 | Aeolia: A Fast and Secure Userspace Interrupt-Based Storage StackabstractPolling-based userspace storage stacks achieve great I/O performance. However, they cannot efficiently and securely share disks and CPUs among multiple tasks. In contrast, interrupt-based kernel stacks inherently suffer from subpar I/O performance but achieve advantages in resource sharing. Chuandong Li 0004, Ran Yi 0004, Zonghao Zhang, Jing Liu 0074, Changwoo Min, Jie Zhang 0048, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Diyu Zhou |
SOSP | 8 |
| 2025 | CortenMM: Efficient Memory Management with Strong Correctness GuaranteesabstractModern memory management systems suffer from poor performance and subtle concurrency bugs, slowing down applications while introducing security vulnerabilities. We observe that both issues stem from the conventional design of memory management systems with two levels of abstraction: a software-level abstraction (e.g., VMA trees in Linux) and a hardware-level abstraction (typically, page tables). This design increases portability but requires correctly and efficiently synchronizing two drastically different and complex data structures, which is generally challenging. Junyang Zhang 0003, Xiangcan Xu, Yonghao Zou, Xinyi Wan 0001, Siyuan Wang 0026, Di Wang 0017, Hao Chen 0023, Lin Huang 0005, Shoumeng Yan, Yuval Tamir, Yingwei Luo, Xiaolin Wang 0001, Huashan Yu, Zhenlin Wang 0003, Hongliang Tian, Diyu Zhou |
SOSP | 15 |
| 2025 | ASTERINAS: A Linux ABI-Compatible, Rust-Based Framekernel OS with a Small and Sound TCB
Yuke Peng, Hongliang Tian, Junyang Zhang 0003, Jinyi Xian, Xiaolin Wang 0001, Chenren Xu, Diyu Zhou, Yingwei Luo, Shoumeng Yan, Yinqian Zhang |
USENIX ATC | 8 |
| 2024 | EKRM: Efficient Key-Value Retrieval Method to Reduce Data Lookup Overhead for Redis
Xiaolin Wang 0001, Diyu Zhou, Liujia Li, Liren Zhu, Zhenlin Wang 0003, Yingwei Luo |
Euro-Par (1) | 2 |
| 2024 | Characterization of Large Language Model Development in the Datacenter
Qinghao Hu 0004, Zhisheng Ye 0002, Zerui Wang, Guoteng Wang, Meng Zhang 0045, Qiaoling Chen, Peng Sun 0006, Dahua Lin, Xiaolin Wang 0001, Yingwei Luo, Yonggang Wen 0001, Tianwei Zhang 0004 |
NSDI | 9 |
| 2024 | Taming Hot Bloat Under Virtualization with HUGESCOPE
Chuandong Li 0004, Sai Sha, Yangqing Zeng, Xiran Yang, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Diyu Zhou |
USENIX ATC | 6 |
| 2024 | Hardware-Software Collaborative Tiered-Memory Management Framework for VirtualizationabstractThe tiered-memory system can effectively expand the memory capacity for virtual machines (VMs). However, virtualization introduces new challenges specifically in enforcing performance isolation, minimizing context switching, and providing resource overcommit. None of the state-of-the-art designs consider virtualization and address these challenges; we observe that a VM with tiered memory incurs up to a 2× slowdown compared to a DRAM-only VM. We propose vTMM , a hardware-software collaborative tiered-memory management framework for virtualization. A key insight in vTMM is to leverage the unique system features in virtualization to meet the above challenges. vTMM automatically determines page hotness and migrates pages between fast and slow memory to achieve better performance. Specially, vTMM optimizes page tracking and migration based on page-modification logging (PML), a hardware-assisted virtualization mechanism, and adaptively distinguishes hot/cold pages through the page “temperature” sorting. vTMM also dynamically adjusts fast memory among multi-VMs on demand by using a memory pool. Further, vTMM tracks huge pages at regular-page granularity in hardware and splits/merges pages in software, realizing hybrid-grained page management and optimization. We implement and evaluate vTMM with single-grained page management on an Intel processor, and the hybrid-grained page management on a Sunway processor with hardware mode supporting hardware/software co-designs. Experiments show that vTMM outperforms existing tiered-memory management designs in virtualization. Sai Sha, Chuandong Li 0004, Xiaolin Wang 0001, Zhenlin Wang 0003, Yingwei Luo |
ACM Trans. Comput. Syst. | 3 |
| 2023 | vTMM: Tiered Memory Management for Virtual MachinesabstractThe memory demand of virtual machines (VMs) is increasing, while the traditional DRAM-only memory system has limited capacity and high power consumption. The tiered memory system can effectively expand the memory capacity and increase the cost efficiency. Virtualization introduces new challenges for memory tiering, specifically enforcing performance isolation, minimizing context switching, and providing resource overcommit. However, none of the state-of-the-art designs consider virtualization and thus address these challenges; we observe that a VM with tiered memory incurs up to a 2× slowdown compared to a DRAM-only VM. Sai Sha, Chuandong Li 0004, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003 |
EuroSys | 4 |
| 2023 | FLORIA: A Fast and Featherlight Approach for Predicting Cache PerformanceabstractThe cache Miss Ratio Curve (MRC) serves a variety of purposes such as cache partitioning, application profiling and code tuning. In this work, we propose a new metric, called cache miss distribution, that describes cache miss behavior over cache sets, for predicting cache MRCs. Based on this metric, we present FLORIA, a software-based, online approach that approximates cache MRCs on commodity systems. By polluting a tunable number of cache lines in some selected cache sets using our designed microbenchmark, the cache miss distribution for the target workload is obtained via hardware performance counters with the support of precise event based sampling (PEBS). A model is developed to predict the MRC of the target workload based on its cache miss distribution. Jun Xiao 0009, Yaocheng Xiang, Xiaolin Wang 0001, Yingwei Luo, Andy D. Pimentel, Zhenlin Wang 0003 |
ICS | 3 |
| 2022 | Original Content Is All You Need! an Empirical Study on Leveraging Answer Summary for WikiHowQA Answer Selection TaskabstractAnswer selection task requires finding appropriate answers to questions from informative but crowdsourced candidates. A key factor impeding its solution by current answer selection approaches is the redundancy and lengthiness issues of crowdsourced answers. Recently, Deng et al. (2020) constructed a new dataset, WikiHowQA, which contains a corresponding reference summary for each original lengthy answer. And their experiments show that leveraging the answer summaries helps to attend the essential information in original lengthy answers and improve the answer selection performance under certain circumstances. However, when given a question and a set of long candidate answers, human beings could effortlessly identify the correct answer without the aid of additional answer summaries since the original answers contain all the information volume that answer summaries contain. In addition, pretrained language models have been shown superior or comparable to human beings on many natural language processing tasks. Motivated by those, we design a series of neural models, either pretraining-based or non-pretraining-based, to check wether the additional answer summaries are helpful for ranking the relevancy degrees of question-answer pairs on WikiHowQA dataset. Extensive automated experiments and hand analysis show that the additional answer summaries are not useful for achieving the best performance. Liang Wen, Houfeng Wang, Yingwei Luo, Xiaolin Wang 0001, Xiaodong Zhang 0022, Zhicong Cheng, Dawei Yin 0001 |
COLING | 5 |
| 2022 | M3: A Multi-View Fusion and Multi-Decoding Network for Multi-Document Reading ComprehensionabstractMulti-document reading comprehension task requires collecting evidences from different documents for answering questions.Previous research works either use the extractive modeling method to naively integrate the scores from different documents on the encoder side or use the generative modeling method to collect the clues from different documents on the decoder side individually.However, any single modeling method cannot make full of the advantages of both.In this work, we propose a novel method that tries to employ a multi-view fusion and multi-decoding mechanism to achieve it.For one thing, our approach leverages question-centered fusion mechanism and cross-attention mechanism to gather finegrained fusion of evidence clues from different documents in the encoder and decoder concurrently.For another, our method simultaneously employs both the extractive decoding approach and the generative decoding method to effectively guide the training process.Compared with existing methods, our method can perform both extractive decoding and generative decoding independently and optionally.Our experiments on two mainstream multi-document reading comprehension datasets (Natural Questions and Triv-iaQA) demonstrate that our method can provide consistent improvements over previous state-of-the-art methods. Liang Wen, Houfeng Wang, Yingwei Luo, Xiaolin Wang 0001 |
EMNLP | 4 |
| 2022 | A Question-Oriented Propagation Network for News Reading ComprehensionabstractMachine reading comprehension of news articles remains to be a challenging task since the lengths of its context documents are long. Such reading comprehension task usually requires document-level language understanding while state-of-the-art, pretrained question answering models can only encode sequences with a predefined length limit. In this paper, we propose a novel Question-Oriented Propagation Network (QOPN) model for such task. Specifically, our proposed QOPN first uses a context encoding module to find local question-related clues. Then, it employs a multi-step reasoning module to aggregate question-focused information for iterative reasoning. The novel design put emphasis on capturing question-related information and allow long-range information integration, which is especially beneficial for long-context reading comprehension task. Experiments on two challenging machine comprehension datasets show that the proposed QOPN significantly outperforms previous state-of-the-art models. Liang Wen, Houfeng Wang, Dehong Ma, Yingwei Luo, Xiaolin Wang 0001, Daiting Shi, Zhicong Cheng, Dawei Yin 0001 |
ICASSP | 6 |
| 2022 | Tear Up the Bubble Boom: Lessons Learned From a Deep Learning Research and Development ClusterabstractWith the proliferation of deep learning, there exists a strong need to efficiently operate GPU clusters for deep learning production in giant AI companies, as well as for research and development (R&D) in small-sized research institutes and universities. Existing works have performed thorough trace analysis on large-scale production-level clusters in giant companies, which discloses the characteristics of deep learning production jobs and motivates the design of scheduling frameworks. However, R&D clusters significantly differ from production-level clusters in both job properties and user behaviors, calling for a different scheduling mechanism. In this paper, we present a detailed workload characterization of an R&D cluster, CloudBrain-I, in a research institute, Peng Cheng Laboratory. After analyzing the fine-grained resource utilization, we discover a severe problem for R&D clusters, resource underutilization, which is especially important in R&D clusters while not characterised by existing works. We further investigate two specific underutilization phenomena and conclude several implications and lessons on R&D cluster scheduling. The traces will be open-sourced to motivate further studies in the community. Zehua Yang, Zhisheng Ye 0002, Tianhao Fu, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Tianwei Zhang 0004 |
ICCD | 7 |
| 2022 | Exploring GNN based program embedding technologies for binary related tasksabstractWith the rapid growth of program scale, program analysis, maintenance and optimization become increasingly diverse and complex. Applying learning-assisted methodologies onto program analysis has attracted ever-increasing attention. However, a large number of program factors including syntax structures, semantics, running platforms and compilation configurations block the effective realization of these methods. To overcome these obstacles, existing works prefer to be on a basis of source code or abstract syntax tree, but unfortunately are sub-optimal for binary-oriented analysis tasks closely related to the compilation process. To this end, we propose a new program analysis approach that aims at solving program-level and procedure-level tasks with one model, by taking advantage of the great power of graph neural networks from the level of binary code. By fusing the semantics of control flow graphs, data flow graphs and call graphs into one model, and embedding instructions and values simultaneously, our method can effectively work around emerging compilation-related problems. By testing the proposed method on two tasks, binary similarity detection and dead store prediction, the results show that our method is able to achieve as high accuracy as 83.25%, and 82.77%. Pengcheng Li 0001, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003 |
ICPC | 4 |
| 2022 | Graph Neural Networks Based Memory Inefficiency Detection Using Selective SamplingabstractProduction software of data centers oftentimes suffers from unnecessary memory inefficiencies caused by inappropriate use of data structures, conservative compiler optimizations, and so forth. Nevertheless, whole-program monitoring tools often incur incredibly high overhead due to fine-grained memory access instrumentation. Consequently, the fine-grained monitoring tools are not viable for long-running, large-scale data center applications due to strict latency criteria (e.g., service-level agreement or SLA). To this end, this work presents a novel learning-aided system, namely Puffin, to identify three kinds of unnecessary memory operations including dead stores, silent loads and silent stores, by applying gated graph neural networks onto fused static and dynamic program semantics with respect to relative positional embedding. To deploy the system in large-scale data centers, this work explores a sampling-based detection infrastructure with high efficacy and negligible overhead. We evaluate Puffin upon the well-known SPEC CPU 2017 benchmark suite for four compilation options. Experimental results show that the proposed method is able to capture the three kinds of memory inefficiencies with as high accuracy as 96% and a reduced checking overhead by$5.66\times$over the state-of-the-art tool. Pengcheng Li 0001, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Xu Liu 0001 |
SC | 4 |
| 2022 | Accelerating Address Translation for Virtualization by Leveraging Hardware ModeabstractThe overhead of memory virtualization remains nontrivial. The traditional shadow paging (TSP) resorts to a shadow page table (SPT) to achieve the native page walk speed, but page table updates require hypervisor interventions. Alternatively, nested paging enables low-overhead page table updates, but utilizes the hardware MMU to perform a long-latency two-dimensional page walk. This paper proposes new memory virtualization solutions based on hardware (machine) mode—the highest CPU privilege level in some architectures like Sunway and RISC-V. A programming interface, running in hardware mode, enables software-implementation of hardware support functions. We first proposeSoftware-based Nested Paging (SNP), which extends the software MMU to perform a two-dimensional page walk in hardware mode. Second, we presentSwift Shadow Paging (SSP), which accomplishes page table synchronization by intercepting TLB flushing in hardware mode. Finally we proposeAccelerated Shadow Paging (ASP)combining SSP and SNP. ASP handles the last-level SPT page faults by walking two-dimensional page tables in hardware mode, which eliminates most hypervisor interventions. This paper systematically compares multiple memory virtualization models by analyzing their designs and evaluating their performance both on a real system and a simulator. The experiments show that the virtualization overhead of ASP is less than 4.5% for all workloads. Sai Sha, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003 |
IEEE Trans. Computers | 4 |
| 2022 | Astraea: A Fair Deep Learning Scheduler for Multi-Tenant GPU ClustersabstractModern GPU clusters are designed to support distributed Deep Learning jobs from multiple tenants concurrently. Each tenant may have varied and dynamic resource demands. Unfortunately, existing GPU schedulers fail to thoroughly consider the fairness among the tenants and jobs, which can result in unbalanced resource allocation and unfair user experience. In this article, we present an efficient solution to provide strong fairness while maintaining high scheduling effectiveness in multi-tenant GPU clusters. First, we introduce a novel Long-Term GPU-time Fairness metric, which can comprehensively evaluate the fairness at both the tenant and job levels, based on both the temporal and spatial impacts of resource allocation. Second, we design a new and practical GPU scheduler,Astraea, to enforce the desired fairness among tenants and jobs. Large-scale evaluations show thatAstraeacan improve tenant fairness by up to 9.42× compared to state-of-the-art schedulers, without sacrificing the average job completion time. Zhisheng Ye 0002, Peng Sun 0006, Wei Gao 0064, Tianwei Zhang 0004, Xiaolin Wang 0001, Shengen Yan, Yingwei Luo |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2021 | GRAPHSPY: Fused Program Semantic Embedding through Graph Neural Networks for Memory EfficiencyabstractProduction software oftentimes suffers from unnecessary memory inefficiencies caused by inappropriate use of data structures, programming abstractions, or conservative compiler optimizations. Unfortunately, existing works often adopt a whole-program fine-grained monitoring method incurring incredibly high overhead. This work proposes a learning-aided approach to identify unnecessary memory operations, by applying several prevalent graph neural network models to extract program semantics with respect to program structure, execution semantics and dynamic states. Results show that the proposed approach captures memory inefficiencies with high accuracy of 95.27% and only around 17% overhead of the state-of-the-art. Pengcheng Li 0001, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003 |
DAC | 4 |
| 2021 | An Edge-Fencing Strategy for Optimizing SSSP Computations on Large-Scale GraphsabstractThe Single-Source Shortest Path (SSSP) problem is to compute the shortest distances in a weighted graph from a source vertex to every other vertex. This paper focuses on parallel efficiency and scalability of SSSP computations on large-scale graphs. We propose an edge-fencing strategy to customize a SSSP algorithm's schedule for every SSSP computation, and devise a path-centric SSSP algorithm with this strategy. This strategy aims at reducing both the relaxed edges and relaxations repeated on each edge. It exploits a few fence values to select the relaxed edges and schedule edge relaxations according to lengths of the created paths. The path-centric algorithm works on a hierarchical graph model, and exploits the edge-fencing strategy to schedule edge relaxations in parallel settings. The hierarchical graph model quantifies the length distribution of shortest paths in large-scale graphs, provides appropriate fence values for every SSSP computation. The algorithm was evaluated on a wide range of synthetic graphs and real-world graphs. The experimental results suggest that our algorithm is efficient and scalable for graphs with skewed degree distributions, and its performance is relatively insensitive to the hierarchical graph model's accuracy. Huashan Yu, Xiaolin Wang 0001, Yingwei Luo |
ICPP | 2 |
| 2021 | Extending MapReduce framework with locality keysabstractThis paper extends the existing MapReduce framework to allow the user programmer to control data locality and reduce communication costs of the shuffle operations in iterative in-memory computation. The programming extension is fully consistent with the style of MapReduce and allows straightforward fast implementation. Xiaolin Wang 0001 |
PPoPP | 3 |
| 2021 | Swift shadow paging (SSP): no write-protection but following TLB flushingabstractVirtualization is a key technique for supporting cloud services and memory virtualization is a major component of virtualization technology. Common memory virtualization mechanisms include shadow paging and hardware-assisted paging. The shadow paging model needs to synchronize shadow/guest page tables whenever there is a guest page table update. In the design of traditional shadow paging (TSP), the guest page table pages are write-protected so the updates can be intercepted by the hypervisor to ensure synchronization. Frequent page table updates cause lots of VM_Exits. Researchers have developed hardware-assisted paging to eliminate this overhead. However, address translation needs to walk a two-dimensional page table. This design significantly increases the overhead of page walk. Sai Sha, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003 |
VEE | 4 |
| 2021 | Penalty- and Locality-aware Memory Allocation in Redis Using Enhanced AETabstractDue to large data volume and low latency requirements of modern web services, the use of an in-memory key-value (KV) cache often becomes an inevitable choice (e.g., Redis and Memcached). The in-memory cache holds hot data, reduces request latency, and alleviates the load on background databases. Inheriting from the traditional hardware cache design, many existing KV cache systems still use recency-based cache replacement algorithms, e.g., least recently used or its approximations. However, the diversity of miss penalty distinguishes a KV cache from a hardware cache. Inadequate consideration of penalty can substantially compromise space utilization and request service time. KV accesses also demonstrate locality, which needs to be coordinated with miss penalty to guide cache management. In this article, we first discuss how to enhance the existing cache model, the Average Eviction Time model, so that it can adapt to modeling a KV cache. After that, we apply the model to Redis and propose pRedis, Penalty- and Locality-aware Memory Allocation in Redis, which synthesizes data locality and miss penalty, in a quantitative manner, to guide memory allocation and replacement in Redis. At the same time, we also explore the diurnal behavior of a KV store and exploit long-term reuse. We replace the original passive eviction mechanism with an automatic dump/load mechanism, to smooth the transition between access peaks and valleys. Our evaluation shows that pRedis effectively reduces the average and tail access latency with minimal time and space overhead. For both real-world and synthetic workloads, our approach delivers an average of 14.0%∼52.3% latency reduction over a state-of-the-art penalty-aware cache management scheme, Hyperbolic Caching (HC), and shows more quantitative predictability of performance. Moreover, we can obtain even lower average latency (1.1%∼5.5%) when dynamically switching policies between pRedis and HC. Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003 |
ACM Trans. Storage | 2 |
| 2020 | Huge Page Friendly Virtualized Memory Management
Sai Sha, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003 |
J. Comput. Sci. Technol. | 4 |
| 2019 | pRedis: Penalty and Locality Aware Memory Allocation in RedisabstractDue to large data volume and low latency requirements of modern web services, the use of in-memory key-value (KV) cache often becomes an inevitable choice (e.g. Redis and Memcached). The in-memory cache holds hot data, reduces request latency, and alleviates the load on background databases. Inheriting from the traditional hardware cache design, many existing KV cache systems still use recency-based cache replacement algorithms, e.g., LRU or its approximations. However, the diversity of miss penalty distinguishes a KV cache from a hardware cache. Inadequate consideration of penalty can substantially compromise space utilization and request service time. KV accesses also demonstrate locality, which needs to be coordinated with miss penalty to guide cache management. Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003 |
SoCC | 3 |
| 2019 | EMBA: Efficient Memory Bandwidth Allocation to Improve Performance on Intel Commodity ProcessorabstractOn multi-core processors, contention on shared resources such as the last level cache (LLC) and memory bandwidth may cause serious performance degradation, which makes efficient resource allocation a critical issue in data centers. Intel recently introduces Memory Bandwidth Allocation (MBA) technology on its Xeon scalable processors, which makes it possible to allocate memory bandwidth in a real system. However, how to make the most of MBA to improve system performance remains an open question. In this work, (1) we formulate a quantitative relationship between a program's performance and its LLC occupancy and memory request rate on commodity processors. (2) Guided by the performance formula, we propose a heuristic bound-aware throttling algorithm to improve system performance and (3) we further develop a hierarchical clustering method to improve the algorithm's efficiency. (4) We implement these algorithms in EMBA, a low-overhead dynamic memory bandwidth scheduling system to improve performance on Intel commodity processors. The results show that, when multiple programs run simultaneously on a multi-core processor whose memory bandwidth is saturated, the programs with high memory bandwidth demand usually use bandwidth inefficiently compared with programs with medium memory bandwidth demand from the perspective of CPU performance. By slightly throttling the former's bandwidth, we can significantly improve the performance of the latter. On average, we improve system performance by 36.9% at the expense of 8.6% bandwidth utilization rate. Yaocheng Xiang, Chencheng Ye 0001, Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003 |
ICPP | 3 |
| 2018 | DCAPS: dynamic cache allocation with partial sharingabstractIn a multicore system, effective management of shared last level cache (LLC), such as hardware/software cache partitioning, has attracted significant research attention. Some eminent progress is that Intel introduced Cache Allocation Technology (CAT) to its commodity processors recently. CAT implements way partitioning and provides software interface to control cache allocation. Unfortunately, CAT can only allocate at way level, which does not scale well for a large thread or program count to serve their various performance goals effectively. This paper proposes Dynamic Cache Allocation with Partial Sharing (DCAPS), a framework that dynamically monitors and predicts a multi-programmed workload's cache demand, and reallocates LLC given a performance target. Further, DCAPS explores partial sharing of a cache partition among programs and thus practically achieves cache allocation at a finer granularity. DCAPS consists of three parts: (1) Online Practical Miss Rate Curve (OPMRC), a low-overhead software technique to predict online miss rate curves (MRCs) of individual programs of a workload; (2) a prediction model that estimates the LLC occupancy of each individual program under any CAT allocation scheme; (3) a simulated annealing algorithm that searches for a near-optimal CAT scheme given a specific performance goal. Our experimental results show that DCAPS is able to optimize for a wide range of performance targets and can scale to a large core count. Yaocheng Xiang, Xiaolin Wang 0001, Zihui Huang, Yingwei Luo, Zhenlin Wang 0003 |
EuroSys | 2 |
| 2018 | Get Out of the Valley: Power-Efficient Address Mapping for GPUsabstractGPU memory systems adopt a multi-dimensional hardware structure to provide the bandwidth necessary to support 100s to 1000s of concurrent threads. On the software side, GPU-compute workloads also use multi-dimensional structures to organize the threads. We observe that these structures can combine unfavorably and create significant resource imbalance in the memory subsystem - causing low performance and poor power-efficiency. The key issue is that it is highly application-dependent which memory address bits exhibit high variability. To solve this problem, we first provide an entropy analysis approach tailored for the highly concurrent memory request behavior in GPU-compute workloads. Our window-based entropy metric captures the information content of each address bit of the memory requests that are likely to co-exist in the memory system at runtime. Using this metric, we find that GPU-compute workloads exhibit entropy valleys distributed throughout the lower order address bits. This indicates that efficient GPU-address mapping schemes need to harvest entropy from broad address-bit ranges and concentrate the entropy into the bits used for channel and bank selection in the memory subsystem. This insight leads us to propose the Page Address Entropy (PAE) mapping scheme which concentrates the entropy of the row, channel and bank bits of the input address into the bank and channel bits of the output address. PAE maps straightforwardly to hardware and can be implemented with a tree of XOR-gates. PAE improves performance by 1.31X and power-efficiency by 1.25X compared to state-of-the-art permutation-based address mapping. Xia Zhao 0004, Magnus Jahre, Zhenlin Wang 0003, Xiaolin Wang 0001, Yingwei Luo, Lieven Eeckhout |
ISCA | 5 |
| 2018 | Fast Miss Ratio Curve Modeling for Storage CacheabstractThe reuse distance (least recently used (LRU) stack distance) is an essential metric for performance prediction and optimization of storage cache. Over the past four decades, there have been steady improvements in the algorithmic efficiency of reuse distance measurement. This progress is accelerating in recent years, both in theory and practical implementation. In this article, we present a kinetic model of LRU cache memory, based on the average eviction time (AET) of the cached data. The AET model enables fast measurement and use of low-cost sampling. It can produce the miss ratio curve in linear time with extremely low space costs. On storage trace benchmarks, AET reduces the time and space costs compared to former techniques. Furthermore, AET is a composable model that can characterize shared cache behavior through sampling and modeling individual programs or traces. Xiameng Hu, Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003, Chen Ding 0001, Chencheng Ye 0001 |
ACM Trans. Storage | 2 |
| 2017 | POSTER: BACM: Barrier-Aware Cache Management for Irregular Memory-Intensive GPGPU WorkloadsabstractGeneral-purpose workloads running on modern graphics processing units (GPGPUs) rely on hardware-based barriers to synchronize warps within a thread block (TB). However, imbalance may exist before reaching a barrier if a GPGPU workload contains irregular memory accesses, i.e., some warps may be critical while others may not. Ideally, cache space should be reserved for the critical warps. Unfortunately, current cache management policies are unaware of the existence of barriers and critical warps, which significantly limits the performance of irregular memory-intensive GPGPU workloads.In this work, we propose Barrier-Aware Cache Management (BACM), which is built on top of two underlying policies: a greedy policy and a friendly policy. The greedy policy does not allow non-critical warps to allocate cache lines in the L1 data cache; only critical warps can. The friendly policy allows non-critical warps to allocate cache lines but only over invalid or lower-priority cache lines. Based on the L1 data cache hit rate of non-critical warps, BACM dynamically chooses between the greedy and friendly policies. By doing so, BACM reserves more cache space to accelerate critical warps, thereby improving overall performance. Experimental results show that BACM achieves an average performance improvement of 24% and 20% compared to the GTO and BAWS policies, respectively. BACM's hardware cost is limited to 96 bytes per streaming multiprocessor. Xia Zhao 0004, Zhibin Yu 0001, Zhenlin Wang 0003, Xiaolin Wang 0001, Yingwei Luo, Lieven Eeckhout |
PACT | 5 |
| 2017 | BACM: Barrier-Aware Cache Management for Irregular Memory-Intensive GPGPU WorkloadsabstractGeneral-purpose workloads running on modern graphics processing units rely on hardware-based barriers to synchronize warps within a thread block (TB). However, imbalance may exist before reaching a barrier if a GPGPU workload contains irregular memory accesses, i.e., some warps may be critical while others may not. Ideally, cache space should be reserved for the critical warps. Unfortunately, current cache management policies are unaware of the existence of barriers and critical warps, which significantly limits the performance of irregular memory-intensive GPGPU workloads. In this paper, we propose Barrier-Aware Cache Management (BACM) which is built on top of two underlying policies: a greedy policy and a friendly policy. The greedy policy does not allow non-critical warps to allocate cache lines in the L1 data cache; only critical warps can. The friendly policy allows non-critical warps to allocate cache lines but only over invalid or lower-priority cache lines. BACM dynamically chooses between the greedy and friendly policies based on the L1 data cache hit rate for the non-critical warps. By doing so, BACM reserves more cache space to accelerate critical warps, thereby improving overall performance. Experimental results show that BACM achieves an average performance improvement of 24% and 20% compared to the GTO and BAWS policies, respectively. BACM's hardware cost is limited to 96 bytes per streaming multiprocessor. Xia Zhao 0004, Zhibin Yu 0001, Zhenlin Wang 0003, Xiaolin Wang 0001, Yingwei Luo, Lieven Eeckhout |
ICCD | 5 |
| 2017 | Evaluating the impacts of hugepage on virtual machines
Xiaolin Wang 0001, Taowei Luo, Zhenlin Wang 0003, Yingwei Luo |
Sci. China Inf. Sci. | 1 |
| 2017 | Optimizing Locality-Aware Memory Management of Key-Value CachesabstractThe in-memory cache system is a performance-critical layer in today's web server architectures. Memcached is one of the most effective, representative, and prevalent among such systems. An important problem is on its memory allocation. The default design does not make the best use of the memory. It is unable to adapt when the demand changes, a problem known as slab calcification. This paper introduces locality-aware memory allocation (LAMA), which addresses the problem by first analyzing the locality of Memcached's requests and then reassigning slabs to minimize the miss ratio or the average response time. By evaluating LAMA using various industry and academic workloads, the paper shows that LAMA outperforms existing techniques in the steady-state performance, the speed of convergence, and the ability to adapt to request pattern changes, and overcome slab calcification. The new solution is close to optimal, achieving over 98 percent of the theoretical potential. Furthermore, LAMA can also be adopted in resource partitioning to guarantee quality-of-service (QoS). Xiameng Hu, Xiaolin Wang 0001, Yingwei Luo, Chen Ding 0001, Song Jiang 0001, Zhenlin Wang 0003 |
IEEE Trans. Computers | 2 |
| 2017 | Optimal Symbiosis and Fair Scheduling in Shared CacheabstractOn multi-core processors, applications are run sharing the cache. This paper presents optimization theory to co-locate applications to minimize cache interference and maximize performance. The theory precisely specifies MRC-based composition, optimization, and correctness conditions. The paper also presents a new technique called footprint symbiosis to obtain the best shared cache performance underfair CPU allocation as well as a new sampling technique which reduces the cost of locality analysis. When sampling and optimization are combined, the paper shows that it takes less than 0.1 second analysis per program to obtain a co-run that is within 1.5 percent of the best possible performance. In an exhaustive evaluation with 12,870 tests, the best prior work improves co-run performance by 56 percent on average. The new optimization improves it by another 29 percent. Without single co-run test, footprint symbiosis is able to choose co-run choices that are just 8 percent slower than the best co-run solutions found with exhaustive testing. Xiameng Hu, Xiaolin Wang 0001, Yechen Li, Yingwei Luo, Chen Ding 0001, Zhenlin Wang 0003 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | Barrier-Aware Warp Scheduling for Throughput ProcessorsabstractParallel GPGPU applications rely on barrier synchronization to align thread block activity. Few prior work has studied and characterized barrier synchronization within a thread block and its impact on performance. In this paper, we find that barriers cause substantial stall cycles in barrier-intensive GPGPU applications although GPGPUs employ lightweight hardware-support barriers. To help investigate the reasons, we define the execution between two adjacent barriers of a thread block as a warp-phase. We find that the execution progress within a warp-phase varies dramatically across warps, which we call warp-phase-divergence. While warp-phase-divergence may result from execution time disparity among warps due to differences in application code or input, and/or shared resource contention, we also pinpoint that warp-phase-divergence may result from warp scheduling. Zhibin Yu 0001, Lieven Eeckhout, Vijay Janapa Reddi, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Cheng-Zhong Xu 0001 |
ICS | 6 |
| 2016 | Kinetic Modeling of Data Eviction in Cache
Xiameng Hu, Xiaolin Wang 0001, Yingwei Luo, Chen Ding 0001, Zhenlin Wang 0003 |
USENIX ATC | 2 |
| 2016 | A survey of cloud resource management for complex engineering applications
Haibao Chen, Song Wu 0001, Hai Jin 0001, Jidong Zhai, Yingwei Luo, Xiaolin Wang 0001 |
Frontiers Comput. Sci. | 7 |
| 2016 | Dynamic Memory Balancing for VirtualizationabstractAllocating memory dynamically for virtual machines (VMs) according to their demands provides significant benefits as well as great challenges. Efficient memory resource management requires knowledge of the memory demands of applications or systems at runtime. A widely proposed approach is to construct a miss ratio curve (MRC) for a VM, which not only summarizes the current working set size (WSS) of the VM but also models the relationship between its performance and the target memory allocation size. Unfortunately, the cost of monitoring and maintaining the MRC structures is nontrivial. This article first introduces a low-cost WSS tracking system with effective optimizations on data structures, as well as an efficient mechanism to decrease the frequency of monitoring. We also propose a Memory Balancer (MEB), which dynamically reallocates guest memory based on the predicted WSS. Our experimental results show that our prediction schemes yield a high accuracy of 95.2% and low overhead of 2%. Furthermore, the overall system throughput can be significantly improved with MEB, which brings a speedup up to 7.4 for two to four VMs and 4.54 for an overcommitted system with 16 VMs. Xiaolin Wang 0001, Fang Hou 0005, Yingwei Luo, Zhenlin Wang 0003 |
ACM Trans. Archit. Code Optim. | 2 |
| 2015 | Improving TLB Performance by Increasing Hugepage RatioabstractLinux supports transparent huge page since 2.6.38. It can automatically map huge pages. But this implementation fails to adjust to page alignment in memory allocation and thus cannot use huge page in some situations. The design is not efficient. Our work aims to increase huge page allocation, so as to improve the utilization ratio of huge page and overall performance. The experimental results show that the optimization almost reaches the upper bound of huge page utilization. This software approach delivers a notable performance improvement for a few benchmarks with moderate overhead in physical memory consumption. Taowei Luo, Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003 |
CCGRID | 2 |
| 2015 | Optimal Footprint Symbiosis in Shared CacheabstractOn multicore processors, applications are run sharing the cache. This paper presents online optimization to collocate applications to minimize cache interference to maximize performance. The paper formulates the optimization problem and solution, presents a new sampling technique for locality analysis and evaluates it in an exhaustive test of 12,870 cases. For locality analysis, previous sampling was two orders of magnitude faster than full-trace analysis. The new sampling reduces the cost by another two orders of magnitude. The best prior work improves co-run performance by 56% on average. The new optimization improves it by another 29%. When sampling and optimization are combined, the paper shows that it takes less than 0.1 second analysis per program to obtain a co-run that is within 1.5% of the best possible performance. Xiaolin Wang 0001, Yechen Li, Yingwei Luo, Xiameng Hu, Jacob Brock, Chen Ding 0001, Zhenlin Wang 0003 |
CCGRID | 1 |
| 2015 | Optimal Cache Partition-SharingabstractWhen a cache is shared by multiple cores, its space may be allocated either by sharing, partitioning, or both. We call the last case partition-sharing. This paper studies partition-sharing as a general solution, and presents a theory an technique for optimizing partition-sharing. We present a theory and a technique to optimize partition sharing. The theory shows that the problem of partition-sharing is reducible to the problem of partitioning. The technique uses dynamic programming to optimize partitioning for overall miss ratio, and for two different kinds of fairness. Finally, the paper evaluates the effect of optimal cache sharing and compares it with conventional solutions for thousands of 4-program co-run groups, with nearly 180 million different ways to share the cache by each co-run group. Optimal partition-sharing is on average 26% better than free-for-all sharing, and 98% better than equal partitioning. We also demonstrate the trade-off between optimal partitioning and fair partitioning. Jacob Brock, Chencheng Ye 0001, Chen Ding 0001, Yechen Li, Xiaolin Wang 0001, Yingwei Luo |
ICPP | 5 |
| 2015 | LAMA: Optimized Locality-aware Memory Allocation for Key-value Cache
Xiameng Hu, Xiaolin Wang 0001, Yechen Li, Yingwei Luo, Chen Ding 0001, Song Jiang 0001, Zhenlin Wang 0003 |
USENIX ATC | 2 |
| 2014 | Performance Metrics and Models for Shared Cache
Chen Ding 0001, Xiaoya Xiang, Bin Bao, Hao Luo 0007, Yingwei Luo, Xiaolin Wang 0001 |
J. Comput. Sci. Technol. | 6 |
| 2013 | Towards Eliminating Memory Virtualization Overhead
Xiaolin Wang 0001, Lingmei Weng, Zhenlin Wang 0003, Yingwei Luo |
APPT | 1 |
| 2013 | Who decides migration? A migration lock mechanism for virtual machinesabstractMigration of virtual machines is an important feature for the management of a virtualized environment. Current strategies for managing migration consider more of resource scheduling and system maintenance, ignoring possible constraints on migration from virtual machines and applications. This paper introduces a migration locking mechanism and its implementation on Xen. The virtual machine migration locking mechanism provides a standard interface for virtual machine users to control migration status of their virtual machines according to their own wishes, such as security, hardware or performance demands and so on. But forced migration locking by the end users or applications could conflict with the management decisions by managers or the virtual machine monitors. How to coordinate different requirements between virtual machine users and managers needs further discussion. Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003 |
CNSM | 1 |
| 2013 | Revisiting memory management on virtualized environmentsabstractWith the evolvement of hardware, 64-bit Central Processing Units (CPUs) and 64-bit Operating Systems (OSs) have dominated the market. This article investigates the performance of virtual memory management of Virtual Machines (VMs) with a large virtual address space in 64-bit OSs, which imposes different pressure on memory virtualization than 32-bit systems. Each of the two conventional memory virtualization approaches, Shadowing Paging (SP) and Hardware-Assisted Paging (HAP), causes different overhead for different applications. Our experiments show that 64-bit applications prefer to run in a VM using SP, while 32-bit applications do not have a uniform preference between SP and HAP. In this article, we trace this inconsistency between 32-bit applications and 64-bit applications to its root cause through a systematic empirical study in Linux systems and discover that the major overhead of SP results from memory management in the 32-bit GNU C library ( glibc ). We propose enhancements to the existing memory management algorithms, which substantially reduce the overhead of SP. Based on the evaluations using SPEC CPU2006, Parsec 2.1, and cloud benchmarks, our results show that SP, with the improved memory allocators, can compete with HAP in almost all cases, in both 64-bit and 32-bit systems. We conclude that without a significant breakthrough in HAP, researchers should pay more attention to SP, which is more flexible and cost effective. Xiaolin Wang 0001, Lingmei Weng, Zhenlin Wang 0003, Yingwei Luo |
ACM Trans. Archit. Code Optim. | 1 |
| 2012 | Live Migrating the Virtual Machine Directly Accessing a Physical NICabstractThis paper proposes a dynamic physical-virtual NIC switching. A virtual machine directly accessing the physical NIC can switch to use the virtual NIC if it is to be migrated. This solution is implemented in network configuration level in Guest OS drivers, so it can be used in different VMMs and with different NICs. The experiments illustrate that our solution does not interrupt the network connections from clients' perspective. After the virtual machine switches to use the virtual NIC, it can be migrated and the switching will not affect the migration performance, including the downtime during the migration. Xiaolin Wang 0001, Yingwei Luo, Xiaoming Li 0001, Zhenlin Wang 0003 |
APSCC | 2 |
| 2012 | A model contract and model integration language for integrating geography models in distributed environmentabstractThis paper presents the model contract and the model integration language, which is used for sharing and reusing geography models in a distributed geography modeling environment. The model contract mainly consists of three parts, the model external description information, the model internal structural information and the model execution flow information. Geographer can structure the model contract to invoke shared models and method to simulate the geography phenomenon using the model integration language. The model integration language is a kind of markup language to express the model contract. Xiaolin Wang 0001, Yingwei Luo |
IGARSS | 1 |
| 2012 | Design model execution engine based on web services for distributed geography modeling environmentabstractThe design of execution engine based on web services for distributed geography modeling environment relies on model contract, model management environment, execution node, data center and control node. The execution procedure of geography model execution engine would be like this: first, geographers construct model contract according to the rule of model contract, and then submit the model contract into control node; Second, the control node parsers the model contract, queries the geography model deployment information, entry point and model input/output information from model management environment, keeps them in memory; Third, control node prepares model data and calls the models which distributed on each execution node for simulation; Finally, the environment returns the result to geographers. The model execution engine adopts multi-thread technology, and could handle multi-job submitted by geographers. Xiaolin Wang 0001, Hongqiang Mao, Yingwei Luo |
IGARSS | 1 |
| 2012 | A Dynamic Cache Partitioning Mechanism under Virtualization EnvironmentabstractCache sharing among multiple computing units on chip is common in today's multi-core processors, and a lot of research has focused on the effective management of shared cache. A software management method called page coloring is commonly used to divide the cache among different applications competing for the same cache entries. Both static and dynamic cache partition mechanism have been implemented in operating system or user-space level. However, few efforts have been made under virtualization environments. Our previous work has provided a static cache partition method based on page coloring in Xen, following that, a dynamic cache partition mechanism called Colored Page Migration (CoPaM) is presented in this paper. Xiaolin Wang 0001, Yechen Li, Yingwei Luo, Xiaoming Li 0001, Zhenlin Wang 0003 |
TrustCom | 1 |
| 2012 | Dynamic cache partitioning based on hot page migration
Xiaolin Wang 0001, Yechen Li, Zhenlin Wang 0003, Yingwei Luo, Xiaoming Li 0001 |
Frontiers Comput. Sci. | 1 |
| 2011 | Model semantic network for massive spatial informationabstractSpatial information is now increasing continuously and is available via the Internet. But facing the abundant spatial information, people recognize many agonizing problems, such as how can spatial information cooperate with each other to solve users' task and how do users know where spatial information is located, what kinds of spatial information can be used, and the way of how to use them. The only way to remedy those problems is to use spatial metadata. Spatial metadata is the description of spatial information and its associated information, which give a semantic expatiation for spatial information. Spatial metadata is now developed as an indispensable and powerful tool for data finding, data exchanging, data managing and data utilization. Currently, there are some standards for spatial metadata such as FGDC and ISO/TC211. But the existing spatial metadata is largely for people to use, and only describes spatial information itself. In this paper, spatial metadata is extended to describe the relations among different spatial information and the distribution of spatial information on the Internet. Based on the extensive spatial metadata, a semantic network model for massive spatial information is presented to support the collaboration among spatial information and make the navigation of spatial information more fast and exact. At last, XML technology is adopted to represent the semantic network for spatial information. Xiaolin Wang 0001, Yingwei Luo |
IGARSS | 1 |
| 2011 | Managing and integrating geography models in distributed environmentabstractFor the current "Model islands" problems which occur in the process of the geography modeling and model sharing, we propose an idea of geography model sharing and reusing method which based on the metadata standard in distributed environment. This article mainly solves how to describe and manage the geography model to achieve the model sharing and reuse. We work on the design of the metadata standard, the integrating standard and the managing environment of the geography model. It is expected to provide a convenient platform for the geographers, so that they can easily reuse any model, which means the heterogeneous models can be well shared. Xiaolin Wang 0001, Yingwei Luo |
IGARSS | 1 |
| 2011 | Sharing and reusing geography models via model execution engineabstractThis article analyzes the sharing and reuse of the geography model in the distributed environment based on the methods of geography model integration and framework. We use the model contract to describe the integration of geography model, and design a model contract language to express. We also design the model execution engine architecture which can execute the model contract language. This is the core function to achieve the sharing and reuse of the geography model in distributed environment. Xiaolin Wang 0001, Hongqiang Mao, Yingwei Luo |
IGARSS | 1 |
| 2011 | Low Cost Working Set Size Tracking
Weiming Zhao, Xinxin Jin, Zhenlin Wang 0003, Xiaolin Wang 0001, Yingwei Luo, Xiaoming Li 0001 |
USENIX ATC | 4 |
| 2011 | Selective hardware/software memory virtualizationabstractAs virtualization becomes a key technique for supporting cloud computing, much effort has been made to reduce virtualization overhead, so a virtualized system can match its native performance. One major overhead is due to memory or page table virtualization. Conventional virtual machines rely on a shadow mechanism to manage page tables, where a shadow page table maintained by the VMM (Virtual Machine Monitor) maps virtual addresses to machine addresses while a guest maintains its own virtual to physical page table. This shadow mechanism will result in expensive VM exits whenever there is a page fault that requires synchronization between the two page tables. To avoid this cost, both Intel and AMD provide hardware assists, EPT (extended page table) and NPT (nested page table), to facilitate address translation. With the hardware assists, the MMU (Memory Management Unit) maintains an ordinary guest page table that translates virtual addresses to guest physical addresses. In addition, the extended page table as provided by EPT translates from guest physical addresses to host physical or machine addresses. NPT works in a similar style. With EPT or NPT, a guest page fault can be handled by the guest itself without triggering VM exits. However, the hardware assists do have their disadvantage compared to the conventional shadow mechanism -- the page walk yields more memory accesses and thus longer latency. Our experimental results show that neither hardware-assisted paging (HAP) nor shadow paging (SP) can be a definite winner. Despite the fact that in over half of the cases, there is no noticeable gap between the two mechanisms, an up to 34% performance gap exists for a few benchmarks. We propose a dynamic switching mechanism that monitors TLB misses and guest page faults on the fly, and dynam-ically switches between the two paging modes. Our experiments show that this new mechanism can match and, sometimes, even beat the better performance of HAP and SP. Xiaolin Wang 0001, Jiarui Zang, Zhenlin Wang 0003, Yingwei Luo, Xiaoming Li 0001 |
VEE | 1 |
| 2011 | A Rule-Based Pretreatment Mechanism for Online Mobile Map DataabstractIn online map service for mobile users, it's necessary to provide different map data according to different application scenarios. For a given scenario, we can preprocess the map data to satisfy the online and mobile requirements. This paper proposes a rule-based pretreatment mechanism for online mobile map data which using a novel mobile map data format Byte-Map. Through the use of pretreatment rules definition and rule engine, we can externalize the decision logic from the business logic in data pretreatment mechanism. We can also separate the data preparer and data pretreatment programmer from each other. In one hand, the data preparer can make definition and configure rules to change the pretreatment procedure, without knowing the program code. In another hand, the data pretreatment programmer can avoid modifying and re-compiling program code for handling map data from different source and type. The rule-based pretreatment mechanism made the online mobile map data pretreatment procedure flexible and extensional. Xiaolin Wang 0001, Yingwei Luo |
VTC Fall | 1 |
| 2010 | The design and implementation of GIS applications based on SOAabstractWith the rapid development of network technologies, SOA (Service Oriented Architecture), which is a methodology of constructing the enterprise distributed software systems, has been widely used nowadays. In this paper we discussed mainly about how to construct a traditional GIS application with the related technologies of SOA. First, we explained the two key problems about module reuse in GIS: encapsulation and combination, and the issues which haven't been addressed till now. Next, we described some simple situations of applying SOA in GIS. At last, we showed PKUMAP, which is a light weighted WebGIS application, using the kernel technologies of SOA, such as SCA, BPEL and so on. Xiaolin Wang 0001, Xiao Pang, Yingwei Luo |
IGARSS | 1 |
| 2010 | Web Service encapsulation of fortran-based geographical modelabstractThis paper take the hydrological model SWAT as an example, explores the encapsulation technology which encapsulate a Fortran source code into a Web Service. In the geographical models, there are a great many of global variable parameters and file reading parameters, according to this situation, we put out a method which makes use of XML file to transform data and file pointer parameters. The encapsulation method is the basis of constructing the model library in the distributed geographical model environment. Xiaolin Wang 0001, Yingwei Luo |
IGARSS | 1 |
| 2010 | Evaluating and Optimizing I/O Virtualization in Kernel-based Virtual Machine (KVM)
Xiaolin Wang 0001, Rongfeng Lai, Zhenlin Wang 0003, Yingwei Luo, Xiaoming Li 0001 |
NPC | 2 |
| 2010 | LBS-p: A LBS Platform Supporting Online Map ServicesabstractThis paper presents LBS-p, a LBS supporting platform, which provides online map service. LBS-p consists of LBS-p Mobile and LBS-p Server. LBS-p Mobile is a Java ME application running on the mobile terminal, which dedicates to the request, management and display of mobile map data. LBS-p Server consists of data pre-processing mechanism, data providing module and LBS-oriented GIS service module. Performance evaluations and a LBS application example show that LBS-p is effective. Xiaolin Wang 0001, Xiao Pang, Yingwei Luo |
VTC Fall | 1 |
| 2010 | Byte-Map: A Novel Mobile Map Format Using Two-Byte CoordinatesabstractThis paper presents Byte-Map, which is a novel mobile map format for mobile map service in mobile devices. Byte-Map is a kind of vector format with different blocks through different levels, and map data of Byte-Map is encapsulated in binary stream. The basic cell of Byte-Map is a block, which is fixed in size of 255 units*255 units according to different coordinates systems and thus the coordinates of all the features in a certain block can be encoded with only two bytes. The experiment shows that Byte-Map has good properties in data volume, data decompressing, data parsing, map displaying and memory cost. Xiaolin Wang 0001, Xiao Pang, Yingwei Luo |
VTC Fall | 1 |
| 2010 | DMM: A dynamic memory mapping model for virtual machines
Haogang Chen 0002, Xiaolin Wang 0001, Zhenlin Wang 0003, Yingwei Luo, Xiaoming Li 0001 |
Sci. China Inf. Sci. | 2 |
| 2010 | Dynamic memory paravirtualization transparent to guest OS
Xiaolin Wang 0001, Yifeng Sun, Yingwei Luo, Zhenlin Wang 0003, Yu Li 0006, Haogang Chen 0002, Xiaoming Li 0001 |
Sci. China Inf. Sci. | 1 |
| 2009 | Fast Live Cloning of Virtual Machine Based on XenabstractVirtual Machine (VM) cloning is to create a replica of a source virtual machine (parent virtual machine); the replica, also called child virtual machine, owns exactly the same executing status as parent virtual machine. Fast live cloning guarantees that, during the period of cloning, the services running on the parent virtual machine observe no performance degradation. There are three important goals for fast live cloning: reducing the total cloning time, minimizing the suspension time of the parent virtual machine, and maximizing resource sharing between the parent virtual machine and the child virtual machine. This paper exploits Copy-on-Write (CoW) mechanism fully to minimize the time of duplicating active main memory and secondary storage of the parent virtual machine. As a result, out experiments show that the total cloning time of a virtual machine on Xen virtual machine monitor can be confined within several hundred milliseconds and the downtime of parent virtual machine is limited to tens of milliseconds, close to VMM’s scheduling interval. Experiments also show that the parent virtual machine highly shares memory with the child virtual machine as long as both virtual machines continue executing the same applications. Yifeng Sun, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Haogang Chen 0002, Xiaoming Li 0001 |
HPCC | 3 |
| 2009 | REMOCA: Hypervisor Remote Disk CacheabstractIn virtual machine (VM) systems, with the increase in the number of VMs and the demands of applications, the main memory is becoming a bottleneck of application performance. To improve paging performance for memory-intensive or I/O-intensive workloads, we propose the hypervisor REMOte disk CAche (REMOCA), which allows a virtual machine to use the memory resources on other physical machines as its cache between its virtual memory and virtual disk devices. The goal of REMOCA is to reduce disk accesses, which is much slower than transferring memory pages over modern interconnect networks. As a result, the average disk I/O latency can be improved. REMOCA is implemented within the hypervisor, by intercepting guest events such as page evictions and disk accesses. This design is transparent to the applications, and is compatible with existing techniques like ballooning and ghost buffer. Moreover, a combination of them can provide a more flexible resource management policy. Our experimental results show that REMOCA can efficiently alleviate the impact of thrashing behavior, and also significantly improve the performance for real-world I/O intensive applications. Haogang Chen 0002, Xiaolin Wang 0001, Zhenlin Wang 0003, Xinxin Jin, Yingwei Luo, Xiaoming Li 0001 |
ISPA | 2 |
| 2009 | A Simple Cache Partitioning Approach in a Virtualized EnvironmentabstractVirtualization is often used in systems for the purpose of offering isolation among applications running in separate virtual machines (VM). Current virtual machine monitors (VMMs) have done a decent job in resource isolation in memory, CPU and I/O devices. However, when looking further into the usage of lower-level shared cache, we notice that one virtual machine’s cache behavior may interfere with another’s due to the uncontrolled cache sharing. In this situation, performance isolation cannot be guaranteed. This paper presents a cache partitioning approach which can be implemented in the VMM. We have implemented this mechanism in Xen VMM using the page coloring technique traditionally applied to the OS. Our VMM-based implementation is fully transparent to the guest OSes. It thus shows the advantages of simplicity and flexibility. Our evaluation shows that our cache partitioning method can work efficiently and improve the performance of co-scheduled applications running within different VMs. In the concurrent workloads selected from the SPEC CPU 2006 benchmarks, our technique achieves a performance improvement by up to 19% for the most sensitive workloads Xinxin Jin, Haogang Chen 0002, Xiaolin Wang 0001, Zhenlin Wang 0003, Yingwei Luo, Xiaoming Li 0001 |
ISPA | 3 |
| 2009 | A Refined Mobile Map Format and Its Application
Yingwei Luo, Xiaolin Wang 0001, Xiao Pang |
SSTD | 2 |
| 2008 | A Rule-Based Event Handling ModelabstractA rule-based event handling model is proposed in the paper, which has a hierarchical architecture with four layers: resource layer, knowledge layer, business layer and representation layer. This model is aimed at distributed information integrating and service scheduling for event handling through rule in knowledge layer. Detail work on knowledge layer is explored, which includes definition of formal business rule, reference to resources in a rule, as well as design and implementation of the rule system. A unified urgent event call-in disposition of city joint emergency response systems is taking as an example of rule-based information integrating and service scheduling. Yingwei Luo, Xiaolin Wang 0001, Xinpeng Liu 0001, Zhou Xing, Xiao Pang |
APSCC | 2 |
| 2008 | Live and incremental whole-system migration of virtual machines using block-bitmapabstractIn this paper, we describe a whole-system live migration scheme, which transfers the whole system run-time state, including CPU state, memory data, and local disk storage, of the virtual machine (VM). To minimize the downtime caused by migrating large disk storage data and keep data integrity and consistency, we propose a three-phase migration (TPM) algorithm. To facilitate the migration back to initial source machine, we use an incremental migration (IM) algorithm to reduce the amount of the data to be migrated. Block-bitmap is used to track all the write accesses to the local disk storage during the migration. Synchronization of the local disk storage in the migration is performed according to the block-bitmap. Experiments show that our algorithms work well even when I/O-intensive workloads are running in the migrated VM. The downtime of the migration is around 100 milliseconds, close to shared-storage migration. Total migration time is greatly reduced using IM. The block-bitmap based synchronization mechanism is simple and effective. Performance overhead of recording all the writes on migrated VM is very low. Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Yifeng Sun, Haogang Chen 0002 |
CLUSTER | 3 |
| 2005 | Spatial Data Channel in a Mobile Navigation System
Yingwei Luo, Guomin Xiong, Xiaolin Wang 0001, Zhuoqun Xu |
ICCSA (2) | 3 |
| 2005 | Ontological Model of Event for Integration of Inter-organization Applications
Wenjun Wang 0002, Yingwei Luo, Xinpeng Liu 0001, Xiaolin Wang 0001, Zhuoqun Xu |
ICCSA (1) | 4 |
| 2005 | XML Approach to Communication Design of WebGIS
Yingwei Luo, Xinpeng Liu 0001, Xiaolin Wang 0001, Zhuoqun Xu |
ICWE | 3 |
| 2005 | The Study and Application of Crime Emergency Ontology Event Model
Wenjun Wang 0002, Yingwei Luo, Xiaolin Wang 0001, Zhuoqun Xu |
KES (4) | 4 |
| 2004 | SOM: A Novel Model for Defining Topological Line-Region Relations
Xiaolin Wang 0001, Yingwei Luo, Zhuoqun Xu |
ICCSA (3) | 1 |
| 2004 | A Component-Based WebGIS Geo-Union
Yingwei Luo, Xiaolin Wang 0001, Zhuoqun Xu |
ICWE | 2 |
| 2004 | GML Based Ubiquitous WebGIS
Yingwei Luo, Baoqi Huang, Jiangong Xu, Xiaolin Wang 0001, Zhuoqun Xu |
SNPD | 4 |
| 2004 | Agent-based Spatial Information Collaboration and Parallel Mechanisms
Yingwei Luo, Xiaolin Wang 0001, Zhuoqun Xu |
SNPD | 2 |
| 2004 | Component-Based WebGIS and Its Spatial Cache Framework
Yingwei Luo, Xiaolin Wang 0001, Zhuoqun Xu |
WAIM | 2 |
| 2003 | Extension of spatial metadata and agent-based spatial Data navigation mechanismabstractFast navigation to distributed spatial data has been the keystone of distributed GIS. In this paper, a hierarchical spatial metadata Database framework is presented based on the existing spatial metadata standard. This framework can efficiently organize the distributed spatial data in network. With the support of the hierarchical spatial metadata Databases, a user-oriented descriptive specification for spatial data requirement is proposed. A map (map-layer) is described as a tuple consisting of four basic elements of . An agent-based searching scheme for spatial query is also introduced, which can satisfy the requirement of fast navigation in terms of spatial query that can often be non-deterministic or with uncertainty. Yingwei Luo, Xiaolin Wang 0001, Zhuoqun Xu |
GIS | 2 |
| 2003 | Extension of spatial metadata for navigating distributed spatial dataabstractNavigating distributed spatial data have been the keystone of distributed GIS. Using spatial metadata is an effective method to locate and access spatial data. In this paper, based on the existing spatial metadata standards, a hierarchical spatial metadata database framework is presented. The framework is extended to describe the distribution of spatial databases, the relations among different spatial databases in network and how to support the fast navigation to spatial data. The framework consists of Global spatial metadata database (Global SMDB), District spatial metadata database (District SMDB) and General spatial metadata database (General SMDB), which can efficiently organize and manage distributed spatial data in network. Spatial data can be navigated quickly with support of the framework. Yingwei Luo, Xiaolin Wang 0001, Zhuoqun Xu |
IGARSS | 2 |