Hui Lu 0001

dblp:65/4062-1 · also Hui Lv 0001 · DBLP profile ↗
← Back
33ranked-venue papers
8as first author
16since 2021 · last 2026
0009-0006-5062-4710ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 5 first-author · 10 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorComputer networks · 2 · 1 first-authorSecurity and privacy · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 AEP: Achieving Hierarchical Fault Tolerance in DSM Through Atomic Execution Protection
abstract
Recent advances in CXL 3.0 have revitalized Distributed Shared Memory (DSM), enabling zero-copy sharing across machine nodes at near-NUMA latency and making shared-memory communication practical for distributed applications. However, DSM memory allocators remain vulnerable to partial failures: client crashes during critical, non-atomic memory operations can corrupt metadata and stall the entire system. Existing solutions either block all clients during recovery or impose heavy runtime overhead due to distributed fault-tolerance protocols. We propose Atomic Execution Protection (AEP), a kernel-assisted fault tolerance mechanism that guarantees atomic completion of user-space critical operations by deferring termination until protected regions finish. AEP tolerates intra-node partial failures without costly distributed coordination. To extend beyond a single node, we further design AEP-DSM, a hierarchical DSM allocator that combines AEP-protected node-level allocators with cross-node coordination. Our Linux AEP prototype integrated with a cross-node DSM consistency protocol (CXL-SHM) achieves near-native local performance and improves throughput by up to 13.3x over state-of-the-art DSM allocators.
Zixuan Wang 0030, Hang Huang, Jia Rao, Hui Lu 0001, Hao Fan 0006, Song Wu 0001, Hai Jin 0001
EuroSys5
2026 Scaling Attention Beyond GPUs for LLM Inference
abstract
Scaling inference for large language models is increasingly constrained by limited GPU memory, primarily due to the expanding intermediate states (KV caches) required for long-context generation and multi-user workloads. Once the KV cache exceeds the capacity of high-bandwidth memory, it must be offloaded to host memory and reloaded on demand, a workflow severely bottlenecked by the CPU–GPU interconnect, typically PCIe. Existing approaches exploiting offload KV caches to CPU memory and selectively reload partial segments for attention computation often underutilize CPU compute resources and suffer from accuracy degradation. We present Beyond, a drop-in runtime that integrates a smart offloading scheme to selectively identify and retain salient KV entries across continuous decoding sessions, together with a hybrid CPU–GPU attention mechanism for scalable inference. Beyond executes dense attention over recent KV entries stored in GPU memory while performing parallel, per-head sparse attention on salient contextual KV entries residing in CPU memory. The outputs are fused efficiently through a log-sum-exp scheme. During the bandwidth-constrained decoding phase, oversized KV caches are processed cooperatively by the aggregated CPU and GPU memory bandwidth, with only minimal PCIe data movement. Experiments across diverse models and workloads demonstrate that Beyond improves scalability, supports longer sequences and larger batch sizes, and outperforms existing sparse attention baselines in both efficiency and accuracy—all on commodity GPU hardware.
Weishu Deng, Peiran Du, Lingfeng Xiang, Chen Zhong 0002, Faraz Ahmed, Lianjie Cao, Puneet Sharma 0001, Song Jiang 0001, Hui Lu 0001, Jia Rao
HPDC11
2026 DLRover-LM: LLM Pre-Training Framework With Thousands of Accelerators in AntGroup
Ziling Huang, Zhengmao Ye, Qingsong Cai, Zelong Huang, Bo Sang, Jian Sha, Tingfeng Lan, Hui Lu 0001, Yuanchun Zhou, MingJie Tang
ICDE9
2025 Can Hardware Outsmart Software in Tiered Memory Management? A CMM-H Case Study
abstract
With the advent of Compute Express Link (CXL), hardware-managed memory tiering has become a reality. In this paper, we investigate Samsung's CXL Memory Module-Hybrid (CMM-H), a CXL Type 3 device integrating DRAM and NAND flash managed by an FPGA-based controller and providing byte-addressable memory interface via the cxl.mem protocol. We perform a detailed evaluation of CMM-H and compare its performance with OS-level and block-level tiering solutions. Our results highlight the performance benefits of CMM-H for cache-hit scenarios and identify key limitations for cache-miss situations, offering insights into the trade-offs involved in adopting hardware-managed memory tiering in emerging CXL-based systems.
Lingfeng Xiang, Lianjie Cao, Faraz Ahmed, Jia Rao, Hui Lu 0001, Puneet Sharma 0001
SYSTOR7
2025 DSA-2LM: A CPU-Free Tiered Memory Architecture with Intel DSA
Ruili Liu, Teng Ma 0006, Yingdi Shan, Zheng Liu 0022, Lingfeng Xiang, Hui Lu 0001, Jia Rao, Kang Chen 0001, Yongwei Wu 0001
USENIX ATC9
2025 LITESHIELD: Secure Containers via Lightweight, Composable Userspace μKernel Services
Kaesi Manakkal, Nathan Daughety, Marcus Pendleton, Hui Lu 0001
USENIX ATC4
2025 mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUs
abstract
Transformer-based large language models (LLMs) have demonstrated outstanding performance across diverse domains, particularly in the emerging pretrain-then-finetune paradigm. LoRA, a parameter-efficient fine-tuning method, is commonly used to adapt a base LLM to multiple downstream tasks. Further, LLM platforms enable developers to fine-tune multiple models and develop various domain-specific applications simultaneously. However, existing model parallelism schemes suffer from high communication overhead and inefficient GPU utilization. In this paper, we present mLoRA, a parallelism-efficient fine-tuning system designed for training multiple LoRA across GPUs and machines. mLoRA introduces a novel LoRA-aware pipeline parallelism scheme that efficiently pipelines LoRA adapters and their distinct fine-tuning stages across GPUs and machines, along with a new LoRA-efficient operator to enhance GPU utilization. Our extensive evaluation shows that mLoRA can significantly reduce average fine-tuning task completion time, e.g., by 30%, compared to state-of-the-art methods like FSDP. More importantly, mLoRA enables simultaneous fine-tuning of larger models, e.g., two Llama-2-13B models on four NVIDIA RTX A6000 48GB GPUs, which is not feasible for FSDP due to high memory requirements. Hence, mLoRA not only increases fine-tuning efficiency but also makes it more accessible on cost-effective GPUs.
Zhengmao Ye, Dengchun Li, Zetao Hu, Tingfeng Lan, Jian Sha, Shicong Zhang, Lei Duan, Jie Zuo, Hui Lu 0001, Yuanchun Zhou, MingJie Tang
Proc. VLDB Endow.9
2024 Nomad: Non-Exclusive Memory Tiering via Transactional Page Migration
Lingfeng Xiang, Weishu Deng, Hui Lu 0001, Jia Rao, Ren Wang 0001
OSDI4
2024 DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the cloud
abstract
Deep learning recommendation models (DLRM) rely on large embedding tables to manage categorical sparse features. Expanding such embedding tables can significantly enhance model performance, but at the cost of increased GPU/CPU/memory usage. Meanwhile, tech companies have built extensive cloud-based services to accelerate training DLRM models at scale. In this paper, we conduct a deep investigation of the DLRM training platforms at AntGroup and reveal two critical challenges: low resource utilization due to suboptimal configurations by users and the tendency to encounter abnormalities due to an unstable cloud environment. To overcome them, we introduce DLRover, an elastic training framework for DLRMs designed to increase resource utilization and handle the instability of a cloud environment. DLRover develops a resource-performance model by considering the unique characteristics of DLRMs and a three-stage heuristic strategy to automatically allocate and dynamically adjust resources for DLRM training jobs for higher resource utilization. Further, DLRover develops multiple mechanisms to ensure efficient and reliable execution of DLRM training jobs. Our extensive evaluation shows that DLRover reduces job completion times by 31%, increases the job completion rate by 6%, enhances CPU usage by 15%, and improves memory utilization by 20%, compared to state-of-the-art resource scheduling frameworks. DLRover has been widely deployed at AntGroup and processes thousands of DLRM training jobs on a daily basis. DLRover is open-sourced and has been adopted by 10+ companies.
Qinlong Wang, Tingfeng Lan, Yinghao Tang, Bo Sang, Ziling Huang, Yiheng Du, Jian Sha, Hui Lu 0001, Yuanchun Zhou, Ke Zhang 0048, MingJie Tang
Proc. VLDB Endow.9
2023 When Caching Systems Meet Emerging Storage Devices: A Case Study
abstract
Block-layer caching systems improve the I/O performance by using hybrid storage devices; the advent of fast, byte-addressable storage enables caching systems to further leverage new storage tiers (e.g., with persistent memory as the cache device and SSD as the backend device) to achieve better caching performance. However, the new storage devices also challenge the design and implementation of existing block-based caching systems. This paper conducts a comprehensive performance study of a popular caching system, Open CAS, and identifies new, unrevealed software bottlenecks. Our observations and root cause analysis cast light on optimizing the software stack of caching systems to incorporate emerging storage technologies.
Lianjie Cao, Faraz Ahmed, Hui Lu 0001, Puneet Sharma 0001
HotStorage4
2023 Accelerating Packet Processing in Container Overlay Networks via Packet-level Parallelism
abstract
Overlay networks serve as the de facto network virtualization technique for providing connectivity among distributed containers. Despite the flexibility in building customized private container networks, overlay networks incur significant performance loss compared to physical networks (i.e., the native). The culprit lies in the inclusion of multiple network processing stages in overlay networks, which prolongs the network processing path and overloads CPU cores. In this paper, we propose mFlow, a novel packet steering approach to parallelize the in-kernel data path of network flows. mFlow exploits packet-level parallelism in the kernel network stack by splitting the packets of the same flow into multiple micro-flows, which can be processed in parallel on multiple cores. mFlow devises new, generic mechanisms for flow splitting while preserving in-order packet delivery with little overhead. Our evaluation with both micro-benchmarks and real-world applications demonstrates the effectiveness of mFlow, with significantly improved performance – e.g., by 81% in TCP throughput and 139% in UDP compared to vanilla overlay networks. mFlow even achieved higher TCP throughput than the native (e.g., 29.8 vs. 26.6 Gbps).
Jiaxin Lei, Manish Munikar, Hui Lu 0001, Jia Rao
IPDPS3
2023 PVM: Efficient Shadow Paging for Deploying Secure Containers in Cloud-native Environment
abstract
In cloud-native environments, containers are often deployed within lightweight virtual machines (VMs) to ensure strong security isolation and privacy protection. With the growing demand for customized cloud services, third-party vendors are turning to infrastructure-as-a-service (IaaS) cloud providers to build their own cloud-native platforms, necessitating the need to run a VM or a guest that hosts containers inside another VM instance leased from an IaaS cloud. State-of-the-art nested virtualization in the x86 architecture relies heavily on the host hypervisor to expose hardware virtualization support to the guest hypervisor, not only complicating cloud management but also raising concerns about an increased attack surface at the host hypervisor.
Hang Huang, Jiangshan Lai, Jia Rao, Hui Lu 0001, Wenlong Hou, Zhengyu He, Weidong Han 0003, Tao Ma 0006, Song Wu 0001
SOSP4
2023 P2CACHE: Exploring Tiered Memory for In-Kernel File Systems Caching
Lingfeng Xiang, Jia Rao, Hui Lu 0001
USENIX ATC4
2022 Prism: Streamlined Packet Processing for Containers with Flow Prioritization
abstract
Advanced high-speed network cards have made packet processing in host operating systems a major performance bottleneck. The kernel network stack gives rise to various sources of overheads that limit the throughput and lengthen the per-packet processing latency. The problem is further exacerbated for short-lived, latency-sensitive network flows such as control packets, online gaming, database requests, etc. — in a highly utilized system, especially in virtualized (containerized) cloud environments, short flows can experience excessively long in-kernel queuing delays. As a consequence, recent research works propose to bypass the kernel network stack to enable lightweight, custom userspace network stacks for improved performance, but at a heavy cost of compatibility and security. In this paper, we take a different approach: We first analyze various sources of inefficiencies in the kernel network stack and propose ways to mitigate them without compromising systems compatibility, security, or flexibility. Further, we propose Prism, a novel mechanism in the kernel network stack to differentiate incoming packets based on their performance requirements and streamline the processing stages of multi-stage packet processing pipelines (e.g., in container overlay networks). Our evaluation demonstrates that Prism can significantly improve the latency of high-priority flows in container overly networks in the presence of heavy low-priority background traffic.
Manish Munikar, Jiaxin Lei, Hui Lu 0001, Jia Rao
ICDCS3
2021 Parallelizing packet processing in container overlay networks
abstract
Container networking, which provides connectivity among containers on multiple hosts, is crucial to building and scaling container-based microservices. While overlay networks are widely adopted in production systems, they cause significant performance degradation in both throughput and latency compared to physical networks. This paper seeks to understand the bottlenecks of in-kernel networking when running container overlay networks. Through profiling and code analysis, we find that a prolonged data path, due to packet transformation in overlay networks, is the culprit of performance loss. Furthermore, existing scaling techniques in the Linux network stack are ineffective for parallelizing the prolonged data path of a single network flow.
Jiaxin Lei, Manish Munikar, Kun Suo, Hui Lu 0001, Jia Rao
EuroSys4
2021 FlashCube: Fast Provisioning of Serverless Functions with Streamlined Container Runtimes
abstract
Fast provisioning of serverless functions is salient for serverless platforms. Though lightweight sandboxes (e.g., containers) enclose only necessary files and libraries, a cold launch still requires up to a few seconds to complete. Such slow provisioning prolongs the response time of serverless functions and negatively impacts users' experiences. This paper analyzes the main reasons for such slowdown and introduces an effective containerization framework, FlashCube. Instead of building a container from scratch, FlashCube quickly and efficiently assembles it through a group of pre-created general container parts (e.g., namespaces, cgroups, and language runtimes). In addition, FlashCube's user-space implementation makes it easily applicable to existing commodity serverless platforms. Our preliminary evaluation demonstrates that FlashCube can quickly provision containerized functions in less than 10 ms (vs. ~400 ms using Docker containers).
Kao-Feng Hsieh, Seunghee Shin, Hui Lu 0001
PLOS@SOSP5
2020 FedMax: Enabling a Highly-Efficient Federated Learning Framework
abstract
IoT devices produce a wealth of data desired for learning models to empower more intelligent applications. However, such data is often privacy sensitive making data owners reluctant upload their data to a central server for learning purposes. Federated learning provides a promising privacy-preserving learning approach, which decouples the model training from the need of accessing to the sensitive data. However, realizing a deployed, dependable federated learning system faces critical challenges, such as frequent dropouts of learning workers, heterogeneity of workers computation, and limited communication. In this paper, we focus on the systems aspects to advance federated learning and contribute a highly efficient and reliable distributed federated learning framework, FedMax, aiming to tackle these challenges. In designing FedMax, we contribute new techniques in light of the properties of a real federated learning setting, including a relaxed synchronization communication scheme and a similarity-based worker selection approach. We have implemented a prototype of FedMax and evaluated FedMax upon multiple popular machine learning models and datasets, showing that FedMax significantly increases the robustness of a federated learning system, speeds up the convergence rate by 25%, and increases the system efficiency by 50%, in comparison with state-of-the-art approaches.
Haohang Xu, Jin Li 0057, Hongkai Xiong, Hui Lu 0001
CLOUD4
2020 Baoverlay: a block-accessible overlay file system for fast and efficient container storage
abstract
Container storage commonly relies on overlay file systems to interpose read-only container images upon backing file systems. While being transparent to and compatible with most existing backing file systems, the overlay file-system approach imposes nontrivial I/O overhead to containerized applications, especially for writes: To write a file originating from a read-only container image, the whole file will be copied to a separate, writable storage layer, resulting in long write latency and inefficient use of container storage. In this paper, we present BAOverlay, a lightweight, block-accessible overlay file system: Equipped with a new block-accessibility attribute, BAOverlay not only exploits the benefit of using an asynchronous copy-on-write mechanism for fast file updates but also enables a new file format for efficient use of container storage space. We have developed a prototype of BAOverlay upon Linux Ext4. Our evaluation with both micro-benchmarks and real-world applications demonstrates the effectiveness of BAOverlay with improved write performance and on-demand container storage usage.
Jiaxin Lei, Seunghee Shin, Hui Lu 0001
SoCC4
2020 PLASMA: programmable elasticity for stateful cloud computing applications
abstract
Developers are always on the lookout for simple solutions to manage their applications on cloud platforms. Major cloud providers have already been offering automatic elasticity management solutions (e.g., AWS Lambda, Azure durable function) to users. However, many cloud applications are stateful --- while executing, functions need to share their state with others. Providing elasticity for such stateful functions is much more challenging, as a deployment/elasticity decision for a stateful entity can strongly affect others in ways which are hard to predict without any application knowledge. Existing solutions either only support stateless applications (e.g., AWS Lambda) or only provide limited elasticity management (e.g., Azure durable function) to stateful applications.
Bo Sang, Pierre-Louis Roman, Patrick Eugster, Hui Lu 0001, Srivatsan Ravi, Gustavo Petri
EuroSys4
2020 SDN-based Order-aware Live Migration of Virtual Machines
abstract
Live migration is a key technique to transfer virtual machines (VMs) from one machine to another. Often multiple VMs need to be migrated in response to events such as server maintenance, load balancing, and impending failures. However, VM migration is a resource intensive operation that pressures the CPU, memory, and network resources of the source and destination hosts as well as intermediate network links. The live migration mechanism ends up contending for finite resources with the VMs that it needs to migrate, which prolongs the total migration time and worsens the performance of applications running inside the VMs. In this paper, we propose SOLive, a new approach to reduce resource contention between the migration process and the VMs being migrated. First, by considering the nature of VM workloads, SOLive manages the order in which multiple VMs are migrated to significantly reduce the total mi-gration time. Secondly, to reduce the network contention between the migration process and the VMs, SOLive uses a combination of software-defined networking-based resource reservation and scatter gather-based VM migration to quickly deprovision the source host. A prototype implementation of our approach in KVM/QEMU platform shows that SOLive quickly evicts VMs from the source host with low impact on VMs' performance.
Dinuni K. Fernando, Hui Lu 0001
INFOCOM3
2019 ShadeNF: Testing Online Network Functions
abstract
The correct implementation of network policies for "in-production" network functions is critical, as it determines the security, availability and performance of a production network. Usually, conducting network testing for these network functions in a live production environment is attractive, as the production environment captures the most exact, realistic dynamic state and vulnerabilities of the system under test. However, doing so also brings potential risks of impacting or even damaging the production system. To address this tension, we present ShadeNF, a novel online platform for testing in-cloud network functions in a production-like environment, without disrupting the real production system. ShadeNF enables such a production-like environment with an exact live clone of production network functions and real production traffic as the test traffic. In designing and implementing ShadeNF, we address several key challenges and contribute new techniques in supporting such a testing platform, including an SDN-based live, consistent snapshot approach, a new programmable forwarding plane, and a scaled test traffic generator. We implement a ShadeNF prototype upon OpenStack and demonstrate that ShadeNF successfully captures the dynamics of production systems, and effectively localizes a range of policy violations in SDN/NFV systems.
Hui Lu 0001, Abhinav Srivastava
IC2E1
2019 Fast and live hypervisor replacement
abstract
Hypervisors are increasingly complex and must be often updated for applying security patches, bug fixes, and feature upgrades. However, in a virtualized cloud infrastructure, updates to an operational hypervisor can be highly disruptive. Before being updated, virtual machines (VMs) running on a hypervisor must be either migrated away or shut down, resulting in downtime, performance loss, and network overhead. We present a new technique, called HyperFresh, to transparently replace a hypervisor with a new updated instance without disrupting any running VMs. A thin shim layer, called the hyperplexor, performs live hypervisor replacement by remapping guest memory to a new updated hypervisor on the same machine. The hyperplexor leverages nested virtualization for hypervisor replacement while minimizing nesting overheads during normal execution. We present a prototype implementation of the hyperplexor on the KVM/QEMU platform that can perform live hypervisor replacement within 10ms. We also demonstrate how a hyperplexor-based approach can used for sub-second relocation of containers for live OS replacement.
Spoorti Doddamani, Piush K. Sinha, Hui Lu 0001, Tsu-Hsiang K. Cheng, Hardik Bagdi, Kartik Gopalan
VEE3
2018 ShadeNF: A Platform for Online Network Function Verification
abstract
The correct implementation of network policies (e.g., routing, NAT, VPNs, load balancing, and IDS/IPS) for underlying network functions is critical, as it determines the security, availability and performance of a production network. However, it is notoriously known that making sure network policies are correctly implemented is challenging, even for basic reachability policies. This becomes more challenging in cloud environments featured with SDN-enabled NFV, where multiple tenants are hosted with richer in-network services in the form of chained, virtualized network functions with dynamic, customized network policies.
Hui Lu 0001, Abhinav Srivastava, Cong Xu 0010
SoCC2
2016 BASS: Improving I/O Performance for Cloud Block Storage via Byte-Addressable Storage Stack
abstract
In an Infrastructure-as-a-Service cloud, cloud block storage offers conventional, block-level storage resources via a storage area network. However, compared to local storage, this multilayered cloud storage model imposes considerable I/O overheads due to much longer I/O path in the virtualized cloud. In this paper, we propose a novel byte-addressable storage stack, BASS, to bridge the addressability gap between the storage and network stacks in cloud, and in return boost I/O performance for cloud block storage. Equipped with byte-addressability, BASS not only avails the benefits of using variable-length I/O requests that avoid unnecessary data transfer, but also enables a highly efficient non-blocking approach that eliminates the blocking of write processes. We have developed a generic prototype of BASS based on Linux storage stack, which is applicable to traditional VMs, lightweight containers and physical machines. Our extensive evaluation with micro-benchmarks, I/O traces and real-world applications demonstrates the effectiveness of BASS, with significantly improved I/O performance and reduced storage network usage.
Hui Lu 0001, Brendan Saltaformaggio, Cong Xu 0010, Umesh Bellur, Dongyan Xu
SoCC1
2016 StorM: Enabling Tenant-Defined Cloud Storage Middle-Box Services
abstract
In an Infrastructure-as-a-Service cloud, tenants rely on the cloud provider to provide "value-added" services such as data security and reliability. However, this provider-controlled service model is less flexible and cannot be customized to meet individual tenants' needs. In this paper, we present StorM, a novel middle-box service platform that allows each tenant to deploy tenant-specific security and reliability services -- in virtualized middle-boxes -- for their cloud data. With such middle-boxes, StorM divides the responsibilities of service creation between tenants and the provider by allowing tenants to customize their own cloud data polices and the provider to offer corresponding infrastructural support. In developing StorM, we address key challenges including network splicing, platform efficiency, and semantic gap. We implement a StorM prototype on top of OpenStack and demonstrate three tenant-defined security/reliability middle-box services, with low performance overhead (<; 10%).
Hui Lu 0001, Abhinav Srivastava, Brendan Saltaformaggio, Dongyan Xu
DSN1
2015 vHaul: Towards Optimal Scheduling of Live Multi-VM Migration for Multi-tier Applications
abstract
Live virtual machine (VM) migration enables seamless movement of an online server from one location to another to achieve failure recovery, load balancing, and system maintenance. Beyond single VM migration, a multi-tier application involves a group of correlated VMs and its live migration will require careful scheduling of the migrations of the member VMs. Our observations from extensive experiments using a variety of multi-tier applications suggest that, in a dedicated data center with dedicated migration links, different migration strategies result in distinct performance impacts on a multi-tier application. The root cause of the problem is the inter-dependence between functional components of a multitier application. We leverage these observations in vHaul, a system that coordinates multi-VM migration to approximate the optimal scheduling. Our evaluation of a vHaul prototype on Xen suggests that vHaul yields the optimal multi-VM live migration schedules. Further, our application-level evaluation using Apache Olio, a web 2.0 cloud application, shows that the optimal migration schedule produced by vHaul outperforms the worst-case schedule by 43% in application throughput. Moreover, the optimal schedule significantly reduces service latency during migration by up to 70%.
Hui Lu 0001, Cong Xu 0010, Ramana Rao Kompella, Dongyan Xu
CLOUD1
2015 vFair: latency-aware fair storage scheduling via per-IO cost-based differentiation
abstract
In virtualized data centers, multiple VMs are consolidated to access a shared storage system. Effective storage resource management, however, turns out to be challenging, as VM workloads exhibit various IO patterns and diverse loads. To multiplex the underlying hardware resources among VMs, providing fairness and isolation while maintaining high resource utilization becomes imperative for effective storage resource management. Existing schedulers such as Linux CFQ or SFQ can provide some fairness, but it has been observed that synchronous IO tends to lose fair shares significantly when competing with aggressive VMs.
Hui Lu 0001, Brendan Saltaformaggio, Ramana Rao Kompella, Dongyan Xu
SoCC1
2013 vTurbo: Accelerating Virtual Machine I/O Processing Using Designated Turbo-Sliced Core
Cong Xu 0010, Sahan Gamage, Hui Lu 0001, Ramana Rao Kompella, Dongyan Xu
USENIX ATC3
2012 Virtualization challenges: a view from server consolidation perspective
abstract
Server consolidation, by running multiple virtual machines on top of a single platform with virtualization, provides an efficient solu-tion to parallelism and utilization of modern multi-core processors system. However, the performance and scalability of server con-solidation solution on modern massive advanced server is not well addressed. In this paper, we conduct a comprehensive study of Xen per-formance and scalability characterization running SPECvirt_sc2010, and identify that large memory and cache footprint, due to the unnecessary high frequent context switch, introduce additional challenges to the system performance and scalability. We propose two optimizations (dynamically-allocable tasklets and context-switch rate controller) to improve the performance. The results show the improved memory and cache efficiency with a reduction of the overall CPI, resulting in an improvement of server consolidation capability by 15% in SPECvirt_sc2010. In the meantime, our optimization achieves an up to 50% acceleration of service response, which greatly improves the QoS of Xen virtualization solution.
Hui Lu 0001, Yaozu Dong, Jiangang Duan, Kevin Tian
VEE1
2009 Graph Matching Based Side Information Generation for Distributed Multi-View Video Coding
abstract
In this paper, we adopt constrained relaxation for distributed multi-view video coding (DMVC). The novel framework integrates the graph-based segmentation and matching to generate inter-view correlated side information without knowing the camera parameters. Moreover, graph-based representations of multi-view images are incorporated to form more distinctive feature constraints. The sparse data as a good hypothesis space aim for a best matching optimization of inter-view side information with compact syndromes, from inferred relaxed coset. The plausible filling-in from a priori feature constraints between neighboring views could reinforce a promising compensation to inter-view side information generation for joint multi-view decoding. In order to find distinctive feature matching with a more stable approximation, PCA-SIFT and TPS (thin plate spline) are adopted to reduce the dimension of SIFT descriptors and construct a more accurate inter-view motion model. The experimental results validate the high estimation precision and the rate-distortion improvements.
Hui Lu 0001, Hongkai Xiong, Li Song 0001, Zhihai He, Tsuhan Chen
ICC1
2009 A Novel Algorithm for Non-dominated Hypervolume-based Multiobjective Optimization
abstract
Hypervolume indicator is a commonly accepted quality measure to assess the set of non-dominated solutions obtained by an evolutionary multiobjective optimization algorithm. Recently, an emerging trend in the design of evolutionary multiobjective optimization algorithms is to directly optimize a quality indicator. In this paper, we propose a hypervolume-based evolutionary algorithm for multiobjective optimization. There are two main contributions of our approach, on one hand, a unique fitness assignment strategy is proposed, on the other hand, we design a slicing based method to calculate the exclusive hypervolume of each individual for environmental selection. From an extensive comparative study with three other MOEAs on a number of two and three objective test problems, it is observed that the proposed algorithm has good performance in convergence and distribution.
Ke Li 0001, Jinhua Zheng, Miqing Li, Hui Lu 0001
SMC5
2008 Side information generation with constrained relaxation for distributed multi-view video coding
abstract
Apart from the existing temporal and inter-view interpolation technique in distributed multi-view video coding (DMVC), this paper is dedicated to not only investigating multiple side information implication at the decoder, but also preserving the constrained relaxation with high-level features matching. We present a novel feature-based Wyner-Ziv coding framework (FWZC) for DMVC, which devotes scale-invariant features extraction and matching to generate inter-view correlated side information without knowing the camera parameters and has a more significant improvement rate-distortion performance. The scale-invariant local features are identified as the most resistant to image deformations and affine distortion between different views of an object or scene. The plausible filling-in from a priori distinctive feature constraints between neighboring views could make a promising compensation to inter-view side information generation for joint multi-view decoding. The experimental results show high precision for objects with high motion and quite significant improvement in the rate-distortion performance.
Hui Lu 0001, Hongkai Xiong, Zhihai He
ISCAS1
2007 A Source-Driven Error Recovery Scheme using Wyner-Ziv Coding
abstract
In this paper, we propose an error recovery scheme to cope with the frame loss problem in large end-to-end delay scenario. Because Wyner-Ziv coding can produce deterministic output with nondeterministic inputs, we adopt the decoder's error-corrupted reference picture as side information to exploit its correlation with source picture, and thus improve the coding efficiency. To prevent retransmission request, we propose a feasible encoder-driven rate estimation scheme by only storing MVs and part of the error pattern at encoder. Furthermore, a new Laplacian parameter computing method is proposed based on discrete Laplacian PDF. The experimental results show that the proposed estimation schemes have quite high precision, and the error recovery scheme outperforms INTRA refresh scheme up to 5 dB.
Hongkai Xiong, Songyu Yu, Hui Lu 0001
ICME4