Qi Zhang 0009

dblp:52/323-9 · DBLP profile ↗
← Back
44ranked-venue papers
9as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 4 first-author · 5 since 2021Systems, architecture and hardware · 12 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Computer networks · 2Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Multi-channel next point-of-interest recommendation with user multi-dimensional preferences
Qi Zhang 0009, Wei Zhou 0028, Junhao Wen 0001
Expert Syst. Appl.1
2025 Disentangling Long- and Short-term Preferences Based on Latent Commonalities Enhancement for Next Point-of-Interest Recommendation
Qi Zhang 0009, Wei Zhou 0028, Junhao Wen 0001
Expert Syst. Appl.3
2024 Deep Learning Service for Efficient Data Distribution Aware Sorting
abstract
In this paper, we present a neural network-enabled data distribution aware sorting method, coined as NN-sort. Our approach explores the potential of developing deep learning techniques to speed up large-scale sort operations, enabling data distribution aware sorting as a deep learning service. Compared to traditional pairwise comparison-based sorting algorithms, which sort data elements by performing pairwise operations, NN-sort leverages the neural network model to learn the data distribution and uses it to map large-scale data elements into ordered ones. Our experiments demonstrate the significant advantage of using NN-sort. Measurements on both synthetic and real-world datasets show that NN-sort yields 2.18× to 10× performance improvement over traditional sorting algorithms.
Xiaoke Zhu, Qi Zhang 0009, Wei Zhou 0011, Ling Liu 0001
IEEE Big Data2
2024 Multi-Grained Trace Collection, Analysis, and Management of Diverse Container Images
abstract
Container technology is getting popular in cloud environments due to its lightweight feature and convenient deployment. Container Registry plays a critical role in container-based clouds, as many container startups involve downloading layer-structured container images from Container Registry. However, Container Registry is struggling to efficiently manage images (i.e., transfer and store) with the emergence of diverse services and new image formats. The reason is that Container Registry manages images uniformly at layer granularity. On the one hand, such uniform layer-level management probably cannot fit the various requirements of different kinds of containerized services well. On the other hand, new image formats organizing data in blocks or files cannot benefit from such uniform layer-level image management. In this paper, we perform the first analysis of image traces at multiple granularities (i.e., image-, layer-, and file-level) for various services and provide an in-depth comparison of different image formats. The traces were collected from a production-level Container Registry, amounting to 24 million requests and involving more than 184 TB of transferred data. We provide a number of valuable insights, including request patterns of services, file-level access patterns, and bottlenecks associated with different image formats. Based on these insights, we propose two optimizations to improve image transfer. Both the traces and toolkit for trace collection will be open-sourced.
Qi Zhang 0009, Hao Fan 0006, Song Wu 0001, Chen Yu 0003, Hai Jin 0001
IEEE Trans. Computers2
2023 Towards Saving Blockchain Fees via Secure and Cost-Effective Batching of Smart-Contract Invocations
abstract
This paper presentsiBatch, a middleware system running on top of an operational Ethereum network to enable secure batching of smart-contract invocations against an untrusted relay server off-chain.iBatchdoes so at a low overhead by validating the server's batched invocations in smart contracts without additional states of user nonces. TheiBatchmechanism supports a variety of policies, ranging from conservative to aggressive batching, and can be configured adaptively to the current workloads.iBatchautomatically rewrites smart contracts to integrate with legacy applications and support large-scale deployment. We built an evaluation platform for fast and cost-accurate transaction replaying and constructed real transaction benchmarks on popular Ethereum applications. With a functional prototype ofiBatch, we conduct extensive cost evaluations, which showsiBatchsaves$14.6\%\sim {}59.1\%$Gas cost per invocation with a moderate 2-minute delay and$19.06\%\sim {}31.52\%$Ether cost per invocation with a delay of$0.26\sim {}1.66$blocks.
Yibo Wang 0006, Kai Li 0017, Yuzhe Tang, Qi Zhang 0009, Xiapu Luo, Ting Chen 0002
IEEE Trans. Software Eng.5
2022 Detecting Layered Bottlenecks in Microservices
abstract
We propose a method to detect both software and hardware bottlenecks in a web service consisting of microservices. A bottleneck is a resource that limits the maximum performance of the entire web service. Bottlenecks often include both software resources such as threads, locks, and channels, and hardware resources such as processors, memories, and disks. Bottlenecks form a layered structure since a single request can utilize multiple software resources and a hardware resource simultaneously. The microservice architecture makes the detection of layered bottlenecks challenging due to the lack of a uniform analysis perspective across languages, libraries, frameworks, and middle-ware.We detect layered bottlenecks in microservices by profiling numbers and status of working threads in each microservice and dependency among microservices via network connections. Our approach can be applied to various programming languages since it relies only on standard debugging tools. Nevertheless, our approach not only detects which microservice is a bottleneck but also enables us to understand why it becomes a bottleneck. This is enabled by a novel visualization method to show layered bottlenecks in microservices at a glance. We demonstrate that our approach successfully detects and visualizes layered bottlenecks in the state-of-the-art microservice benchmarks, DeathStarBench and Acme Air microservices. This enables us to optimize the microservices themselves to achieve a higher throughput per re-source utilization rate compared with simply scaling the number of replicas of microservices.
Tatsushi Inagaki, Yohei Ueda, Moriyoshi Ohara, Sunyanan Choochotkaew, Marcelo Amaral, Scott Trent, Tatsuhiro Chiba, Qi Zhang 0009
CLOUD8
2022 A Comparative Measurement Study of Deep Learning as a Service Framework
abstract
Big data powered Deep Learning (DL) and its applications have blossomed in recent years, fueled by three technological trends: a large amount of digitized data openly accessible, a growing number of DL software frameworks in open source and commercial markets, and a selection of affordable parallel computing hardware devices. However, no single DL framework, to date, dominates in terms of performance and accuracy even for baseline classification tasks on standard datasets, making the selection of a DL framework an overwhelming task. This paper takes a holistic approach to conduct empirical comparison and analysis of four representative DL frameworks with three unique contributions.First, given a selection of CPU-GPU configurations, we show that for a specific DL framework, different configurations of its hyper-parameters may have a significant impact on both performance and accuracy of DL applications.Second, to the best of our knowledge, this study is the first to identify the opportunities for improving the training time performance and the accuracy of DL frameworks by configuring parallel computing libraries and tuning individual and multiple hyper-parameters.Third, we also conduct a comparative measurement study on the resource consumption patterns of four DL frameworks and their performance and accuracy implications, including CPU and memory usage, and their correlations to varying settings of hyper-parameters under different configuration combinations of hardware, parallel computing libraries. We argue that this measurement study provides in-depth empirical comparison and analysis of four representative DL frameworks, and offers practical guidance for service providers to deploying and delivering DL as a Service (DLaaS) and for application developers and DLaaS consumers to select the right DL frameworks for the right DL workloads.
Yanzhao Wu 0001, Ling Liu 0001, Calton Pu, Wenqi Cao, Semih Sahin, Wenqi Wei 0001, Qi Zhang 0009
IEEE Trans. Serv. Comput.7
2021 Knowledge & Learning-based Adaptable System for Sensitive Information Identification and Handling
abstract
Diagnostic data such as logs and memory dumps from production systems are often shared with development teams to do root cause analysis of system crashes. Invariably such diagnostic data contains sensitive information and sharing it can lead to data leaks. To handle this problem we present Knowledge and Learning-based Adaptable System for Sensitive InFormation Identification and Handling (KLASSIFI) which is an end to end system capable of identifying and redacting sensitive information present in diagnostic data. KLASSIFI is highly customizable, allowing it to be used for various different business use cases by simply changing the configuration. KLASSIFI ensures that the output file is useful by retaining the metadata which is used by various debugging tools. Various optimizations have been done to improve the performance of KLASSIFI. Empirical evaluation of KLASSIFI shows that it is able to process large files (128 GB) in 84 minutes and its performance scales linearly with varying factors. This points to practicability of KLASSIFI.
Akshar Kaul, Manish Kesarwani, Hong Min, Qi Zhang 0009
CLOUD4
2021 DLB: Deep Learning Based Load Balancing
abstract
In this paper, we introduce DLB, a Deep Learning based load Balancing mechanism, to effectively address the data skew problem. The key idea of DLB is to replace hash functions in the load balancing mechanisms with deep learning models, which are trained to be able to map different distributions of workloads and data to the servers in a uniformed manner. We implemented DLB and deployed it on a practical Cloud environment using CloudSim. Experimental results using both synthetic and real-world data sets show that compared with traditional hash function based load balancing methods, DLB is able to achieve more balanced mappings, especially when the workload is highly skewed.
Xiaoke Zhu, Qi Zhang 0009, Taining Cheng, Ling Liu 0001, Wei Zhou 0011, Jing He 0012
CLOUD2
2021 iBatch: saving Ethereum fees via secure and cost-effective batching of smart-contract invocations
abstract
This paper presents iBatch, a middleware system running on top of an operational Ethereum network to enable secure batching of smart-contract invocations against an untrusted relay server off-chain. iBatch does so at a low overhead by validating the server's batched invocations in smart contracts without additional states. The iBatch mechanism supports a variety of policies, ranging from conservative to aggressive batching, and can be configured adaptively to the current workloads. iBatch automatically rewrites smart contracts to integrate with legacy applications and support large-scale deployment.
Yibo Wang 0006, Qi Zhang 0009, Kai Li 0017, Yuzhe Tang, Xiapu Luo, Ting Chen 0002
ESEC/SIGSOFT FSE2
2021 BlockHDFS: Blockchain-integrated Hadoop distributed file system for secure provenance traceability
abstract
Hadoop Distributed File System (HDFS) is one of the widely used distributed file systems in big data analysis for frameworks such as Hadoop. HDFS allows one to manage large volumes of data using low-cost commodity hardware. However, vulnerabilities in HDFS can be exploited for nefarious activities. This reinforces the importance of ensuring robust security to facilitate file sharing in Hadoop as well as having a trusted mechanism to check the authenticity of shared files. This is the focus of this paper, where we aim to improve the security of HDFS using a blockchain-enabled approach (hereafter referred to as BlockHDFS). Specifically, the proposed BlockHDFS uses the enterprise-level Hyperledger Fabric platform to capitalize on files' metadata for building trusted data security and traceability in HDFS.
Viraaji Mothukuri, Sai S. Cheerla, Reza M. Parizi, Qi Zhang 0009, Kim-Kwang Raymond Choo
Blockchain Res. Appl.4
2020 Blockchain smart contracts formalization: Approaches and challenges to address vulnerabilities
Amritraj Singh, Reza M. Parizi, Qi Zhang 0009, Kim-Kwang Raymond Choo, Ali Dehghantanha
Comput. Secur.3
2020 SST: Synchronized Spatial-Temporal Trajectory Similarity Search
Weixiong Rao, Chengxi Zhang, Gong Su, Qi Zhang 0009
GeoInformatica5
2020 Sidechain technologies in blockchain networks: An examination and state-of-the-art review
Amritraj Singh, Kelly Click, Reza M. Parizi, Qi Zhang 0009, Ali Dehghantanha, Kim-Kwang Raymond Choo
J. Netw. Comput. Appl.4
2020 GraphMap: scalable iterative graph processing using NoSQL
Sayan Goswami, Ayam Pokhrel, Kisung Lee, Ling Liu 0001, Qi Zhang 0009, Yang Zhou 0001
J. Supercomput.5
2020 Improving Collaborative Filtering with Social Influence over Heterogeneous Information Networks
abstract
The advent of social networks and activity networks affords us an opportunity of utilizing explicit social information and activity information to improve the quality of recommendation in the presence of data sparsity. In this article, we present a social-influence-based collaborative filtering (SICF) framework over heterogeneous information networks with three unique features. First, we integrate different types of entities, links, attributes, and activities from rating networks, social networks, and activity networks into a unified social-influence-based collaborative filtering model through the intra-network and inter-network social influence. Second, we propose three social-influence propagation models to capture three kinds of information propagation within heterogeneous information networks: user-based influence propagation on user rating networks, item-based influence propagation on user-rating activity networks, and term-based influence propagation on user-review activity networks, respectively. We compute three kinds of social-influence-based user similarity scores based on three social-influence propagation models, respectively. Third, a unified social-influence-based CF prediction model is proposed to infer rating tastes by incorporating three kinds of social-influence-based similarity measures with different weighting factors. We design a weight-learning algorithm, SICF, to refine the prediction result by quantifying the contribution of each kind of information propagation to make a good balance between prediction accuracy and data sparsity. Extensive evaluation on real datasets demonstrates that SICF outperforms existing representative collaborative filtering methods.
Yang Zhou 0001, Ling Liu 0001, Kisung Lee, Balaji Palanisamy, Qi Zhang 0009
ACM Trans. Internet Techn.5
2020 An Energy-Efficient SDN Controller Architecture for IoT Networks With Blockchain-Based Security
abstract
Internet of Things (IoT) is a disruptive technology in many aspects of our society, ranging from communications to financial transactions to national security (e.g., Internet of Battlefield / Military Things), and so on. There are long-standing challenges in IoT, such as security, comparability, energy consumption, and heterogeneity of devices. Security and energy aspects play important roles in data transmission across IoT and edge networks, due to limited energy and computing (e.g., processing and storage) resources of networked devices. Whether malicious or accidental, interference with data in an IoT network potentially has real-world consequences. In this article, we explore the potential of integrating blockchain and software-defined networking (SDN) in mitigating some of the challenges. Specifically, we propose a secure and energy-efficient blockchain-enabled architecture of SDN controllers for IoT networks using a cluster structure with a new routing protocol. The architecture uses public and private blockchains for Peer to Peer (P2P) communication between IoT devices and SDN controllers, which eliminates Proof-of-Work (POW), as well as using an efficient authentication method with the distributed trust, making the blockchain suitable for resource-constrained IoT devices. The experimental results indicate that the routing protocol based on the cluster structure has higher throughput, lower delay, and lower energy consumption than EESCFD, SMSN, AODV, AOMDV, and DSDV routing protocols. In other words, our proposed architecture is demonstrated to outperform classic blockchain.
Abbas Yazdinejad, Reza M. Parizi, Ali Dehghantanha, Qi Zhang 0009, Kim-Kwang Raymond Choo
IEEE Trans. Serv. Comput.4
2019 Demystifying Learning Rate Policies for High Accuracy Training of Deep Neural Networks
abstract
Learning Rate (LR) is an important hyper-parameter to tune for effective training of deep neural networks (DNNs). Even for the baseline of a constant learning rate, it is non-trivial to choose a good constant value for training a DNN. Dynamic learning rates involve multi-step tuning of LR values at various stages of the training process and offer high accuracy and fast convergence. However, they are much harder to tune. In this paper, we present a comprehensive study of 13 learning rate functions and their associated LR policies by examining their range parameters, step parameters, and value update parameters. We propose a set of metrics for evaluating and selecting LR policies, including the classification confidence, variance, cost, and robustness, and implement them in LRBench, an LR benchmarking system. LRBench can assist end-users and DNN developers to select good LR policies and avoid bad LR policies for training their DNNs. We tested LRBench on Caffe, an open source deep learning framework, to showcase the tuning optimization of LR policies. Evaluated through extensive experiments, we attempt to demystify the tuning of LR policies by identifying good LR policies with effective LR value ranges and step sizes for LR update schedules.
Yanzhao Wu 0001, Ling Liu 0001, Juhyun Bae, Ka-Ho Chow 0001, Arun Iyengar, Calton Pu, Wenqi Wei 0001, Lei Yu 0002, Qi Zhang 0009
IEEE BigData9
2019 Memory Disaggregation: Research Problems and Opportunities
abstract
Memory usage imbalance has been consistently observed in many virtualized Clouds and production datacenters. Such temporal memory utilization variance is a major root cause for excessive paging and thrashing on virtual servers even though there are sufficient idle memory on the same node or in the Cloud cluster. Memory disaggregation is an emerging research and development endeavor towards addressing these memory usage imbalance problems. This paper first defines and characterizes the concept of memory disaggregation, and discusses the demands and challenges of efficient memory disaggregation in cloud datacenters. It then examines some promising research issues, design choices and directions to overcome some of the challenges posed by memory disaggregation. Specifically, it proposes two major new research challenges and solution directions for enabling elastic, on-demand disaggregated memory orchestration: (1) virtual server memory and node level memory co-design and (2) local memory and remote memory co-design. A brief description of two ongoing research projects is provided for both solution directions. The paper ends with a brief discussion of other advanced and emerging memory and storage technologies and potential opportunities for memory disaggregation.
Ling Liu 0001, Wenqi Cao, Semih Sahin, Qi Zhang 0009, Juhyun Bae, Yanzhao Wu 0001
ICDCS4
2019 A Social Recommendation Algorithm with Trust and Distrust Considering Domain Relevance
Ling Liu 0001, Qi Zhang 0009, Junhao Wen 0001
ICONIP (5)2
2019 An Efficient and Transparent Approach for Adaptive Intra-and Inter-Node Virtual Machine Communication in Virtualized Clouds
abstract
Network I/O workloads are dominating as one of the leading costs for most of the virtualized clouds. One way to improve the inter virtual machine (VM) inefficiency is to build shared memory channels between VMs co-located on the same physical node to by-pass traditional TCP/IP network stack, so that the overhead is reduced by shorter communication path and fewer kernel interactions. However, it is a key challenge for existing work to achieve high performance inter-VM communication while keeping the capability of VM live migration, and most of existing work are neither seamlessly agile in the presence of VM live migration nor transparent to upper users as well as to operating system kernels, which limits the application of current co-location aware shared-memory based approaches. In this paper, we present the design and implementation of XenVMC, an adaptive and transparent inter-VM communication system for high performance network I/O in virtualized clouds. With proposed dynamic co-located VM membership update mechanism, XenVMC is applicable not only to intra-node VM communication, but also to cross-node communication. It also supports adaptive switching between shared-memory based channel and traditional network-based channel in case of VM live migration, with the aid of proposed VM migration perception and handling mechanisms. XenVMC enables efficient data transmission for both TCP and UDP workloads, with multilevel transparency guaranteed. Extensive experiments show that XenVMC achieves better performance for both TCP and UDP workloads with high transparency, compared with both native virtualized environment and representative existing work. Experimental results also show that it is capable of automatically handling VM migration correctly with acceptable latency.
Yi Ren 0008, Renshi Liu, Qi Zhang 0009, Jianbo Guan, Ziqi You, Yusong Tan, Qingbo Wu 0003
ICPADS3
2019 CLEAN: Frequent Pattern-Based Trajectory Spatial-Temporal Compression on Road Networks
abstract
The volume of trajectory data has become tremendously large in recent years. How to efficiently maintain and compute such trajectory data becomes a challenging task. In this paper, we propose a trajectory spatial and temporal compression framework, namely CLEAN. The key of spatial compression is to mine meaningful trajectory frequent patterns on road networks. By treating the mined patterns as dictionary items, we have the chance to encode a long trajectory by shorter paths, thus leading to smaller space cost. Meanwhile, we design an error-bounded temporal compression on top of the identified spatial patterns for much low space cost. Extensive experiments on real trajectory datasets validate that CLEAN significantly outperforms existing state-of-art approaches in terms of both space saving and runtime of trajectory compression.
Qinpei Zhao, Chenxi Zhang 0001, Gong Su, Qi Zhang 0009, Weixiong Rao
MDM5
2019 Lightweight Indexing and Querying Services for Big Spatial Data
abstract
With the widespread use of GPS-equipped smartphones and Internet of Things devices, a huge amount of data with location information is being generated at an unprecedented rate. To gain a deeper insight into such a plethora of spatial data, scientists and engineers are widely using spatial queries for their big data applications. However, because of not only the massive spatial data size but also the complexity of spatial query processing, they are struggling to efficiently process the spatial queries. In this paper, we propose lightweight and scalable indexing and querying services for big spatial data stored in distributed storage systems or graph-based systems. Our spatial services have several advantages over existing approaches. First, our services can be easily applied to existing storage systems or graph-based models without modifying the internal implementation of existing systems/models. Second, our services achieve high pruning power by efficiently selecting only relevant spatial objects based on a simple yet effective filter. Third, our services support a customizable and easy-to-use control of index data size by adjusting the precision of indexed geometries. Lastly, our services support efficient updates of spatial data. Our experimental results using real-world datasets validate the effectiveness and efficiency of our spatial services.
Kisung Lee, Ling Liu 0001, Raghu K. Ganti, Mudhakar Srivatsa, Qi Zhang 0009, Yang Zhou 0001, Qingyang Wang 0001
IEEE Trans. Serv. Comput.5
2018 A Comparative Study of Containers and Virtual Machines in Big Data Environment
abstract
Container technique is gaining increasing attention in recent years and has become an alternative to traditional virtual machines. Some of the primary motivations for the enterprise to adopt the container technology include its conveniency to encapsulate and deploy applications, lightweight operations, as well as efficiency and flexibility in resources sharing. However, there still lacks an in-depth and systematic comparison study on how big data applications, such as Spark jobs, perform between a container environment and a virtual machine environment. In this paper, by running various Spark applications with different configurations, we evaluate the two environments from many interesting aspects, such as how convenient the execution environment can be set up, what are makespans of different workloads running in each setup, how efficient the hardware resources, such as CPU and memory, are utilized, and how well each environment can scale. The results show that compared with virtual machines, containers provide a more easy-to-deploy and scalable environment for big data workloads. The research work in this paper can help practitioners and researchers to make more informed decisions on tuning their cloud environment and configuring the big data applications, so as to achieve better performance and higher resources utilization.
Qi Zhang 0009, Ling Liu 0001, Calton Pu, Qiwei Dou, Liren Wu, Wei Zhou 0011
IEEE CLOUD1
2018 Benchmarking Deep Learning Frameworks: Design Considerations, Metrics and Beyond
abstract
With increasing number of open-source deep learning (DL) software tools made available, benchmarking DL software frameworks and systems is in high demand. This paper presents design considerations, metrics and challenges towards developing an effective benchmark for DL software frameworks and illustrate our observations through a comparative study of three popular DL frameworks: TensorFlow, Caffe, and Torch. First, we show that these deep learning frameworks are optimized with their default configurations settings. However, the default configuration optimized on one specific dataset may not work effectively for other datasets with respect to runtime performance and learning accuracy. Second, the default configuration optimized on a dataset by one DL framework does not work well for another DL framework on the same dataset. Third, experiments show that different DL frameworks exhibit different levels of robustness against adversarial examples. Through this study, we envision that unlike traditional performance-driven benchmarks, benchmarking deep learning software frameworks should take into account of both runtime and accuracy and their latent interaction with hyper-parameters and data-dependent configurations of DL frameworks.
Ling Liu 0001, Yanzhao Wu 0001, Wenqi Wei 0001, Wenqi Cao, Semih Sahin, Qi Zhang 0009
ICDCS6
2018 Efficient Shared Memory Orchestration towards Demand Driven Memory Slicing
abstract
Memory is increasingly becoming a bottleneck for big data and latency-sensitive applications in virtualized systems. Memory efficiency is critical for high-performance execution of virtual machines (VMs). Mechanisms proposed for improving memory utilization often rely on an accurate estimation of VM working set size at runtime, which is difficult under changing workloads. This paper explores opportunities for improving memory efficiency and their impacts on the performance of VM executions. First, we show that if each VM is initialized with an application-specified lower bound memory, then by maintaining a shared memory region across VMs in the presence of temporal memory usage variations on the host, those VMs under high memory pressure can minimize their performance loss by opportunistically and transparently harvesting idle memory on other VMs. Second, we show that by enabling on-demand VM memory allocation and deallocation in the presence of changing workloads, VM performance degradation due to memory swapping can be reduced effectively, compared to the conventional VM configuration scenario, in which all VMs are allocated with the upper-bound of memory requested by their applications. Third, we show that by providing shared memory pipes between co-located VMs, the inter-VM communication can speed up by avoiding unnecessary overhead of communication via the network. We develop MemLego, a lightweight shared memory based system, to achieve all these benefits without requiring any modification to user applications and the OSes. We demonstrate the effectiveness of these opportunities through extensive experiments on unmodified Redis and MemCached. Using MemLego, the throughput of Redis and Memcached improves by up to 4x over the native system without MemLego, up to 2 orders of magnitude when the applications working set size does not fit in memory.
Qi Zhang 0009, Ling Liu 0001, Calton Pu, Wenqi Cao, Semih Sahin
ICDCS1
2017 MemFlex: A Shared Memory Swapper for High Performance VM Execution
abstract
Ballooning is a popular solution for dynamic memory balancing. However, existing solutions may perform poorly in the presence of heavy guest swapping. Furthermore, when the host has sufficient free memory, guest virtual machines (VMs) under memory pressure is not be able to use it in a timely fashion. Even after the guest VM has been recharged with sufficient memory via ballooning, the applications running on the VM are unable to utilize the free memory in guest VM to quickly recover from the severe performance degradation. To address these problems, we present MemFlex, a shared memory swapper for improving guest swapping performance in virtualized environment with three novel features: (1) MemFlex effectively utilizes host idle memory by redirecting the VM swapping traffic to the host-guest shared memory area. (2) MemFlex provides a hybrid memory swapping model, which treats a fast but small shared memory swap partition as the primary swap area whenever it is possible, and smoothly transits to the conventional disk-based VM swapping on demand. (3) Upon ballooned with sufficient VM memory, MemFlex provides a fast swap-in optimization, which enables the VM to proactively swap in the pages from the shared memory using an efficient batch implementation. Instead of relying on costly page faults, this optimization offers just-in-time performance recovery by enabling the memory intensive applications to quickly regain their runtime momentum. Performance evaluation results are presented to demonstrate the effectiveness of MemFlex when compared with existing swapping approaches.
Qi Zhang 0009, Ling Liu 0001, Gong Su, Arun Iyengar
IEEE Trans. Computers1
2016 Tenants Attested Trusted Cloud Service
abstract
Cloud computing has successfully enabled large scale computing to be offered as pay-as-you-go services to many enterprise and individual tenants. However, the trust on public cloud services has been a sensitive issue for both cloud tenants and cloud service providers (CSPs). Tenants tend to worry about losing the total control over their codes and data hosted on remote servers. Public cloud providers often fear that the applications uploaded by their tenants may carry vicious codes, which may cause serious violations of security and privacy on their cloud platforms. This trust issue has slowed down the wide deployment of public clouds and hindered the promises of cloud computing for both CSPs and Cloud consumers. In this paper, we present Ta-TCS, a novel system framework for two-phase tenants attested trust validation and trust management over their remote VMs and cloud service executions. At the CSP end, we build a Minimal Trusted Environment (MTE) in VMM and an Integrity Verification & Report Service (IVRS) hosted in the control domain Dom0. At the tenant end, we deploy an Integrity Configuration and Attestation Service (ICAS) in new framework. With Ta-TCS, tenants can configure and attest the integrity of their services, and Cloud providers can verify codes running on a guest VM by introspection. Tenants can also check whether the basic platform of Dom0 is trusted or not. This two phase trust establishment increases the level of mutual trust between tenants and its CSP. We implement the first prototype system of Ta-TCS on Xen platform, and most of our implementation mechanisms can be deployed to some open-source virtualization platforms such as KVM. Our evaluation results show that Ta-TCS is effective with negligible performance overhead.
Jiangchun Ren, Ling Liu 0001, Qi Zhang 0009, Haihe Ba
CLOUD4
2016 iBalloon: Efficient VM Memory Balancing as a Service
abstract
Dynamic VM memory management via the balloon driver is a common strategy to manage the memory resources of VMs under changing workloads. However, current approaches rely on kernel instrumentation to estimate the VM working set size, which usually result in high run-time overhead. Thus system administrators have to tradeoff between the estimation accuracy and the system performance. This paper presents iBalloon, a light-weight, accurate and transparent prediction based mechanism to enable more customizable and efficient ballooning policies for rebalancing memory resources among VMs. Experiment results from well known benchmarks such as Dacapo and SPECjvm show that iBalloon is able to quickly react to the VM memory demands, provide up to 54% performance speedup for memory intensive applications running in the VMs, while incurring less than 5% CPU overhead on the host machine as well as the VMs.
Qi Zhang 0009, Ling Liu 0001, Jiangchun Ren, Gong Su, Arun Iyengar
ICWS1
2016 Workload Adaptive Shared Memory Management for High Performance Network I/O in Virtualized Cloud
abstract
This paper presents the design and implementation of MemPipe, a dynamic shared memory management system for high performance network I/O among virtual machines (VMs) located on the same host. MemPipe delivers efficient inter-VM communication with three unique features. First, MemPipe employs an inter-VM shared memory pipe to enable high throughput data delivery for both TCP and UDP workloads among co-located VMs. Second, instead of static allocation of shared memories, MemPipe manages its shared memory pipes through a demand driven and proportional memory allocation mechanism, which can dynamically enlarge or shrink the shared memory pipes based on the demand of the workloads in each VM. Third but not the least, MemPipe employs a number of optimizations, such as time-window based streaming partitions and socket buffer redirection, to further optimize its performance. Extensive experiments show that MemPipe improves the throughput of conventional (native) inter VM communication by up to 45 times, reduces the latency by up to 62 percent, and achieves up to 91 percent shared memory utilization.
Qi Zhang 0009, Ling Liu 0001
IEEE Trans. Computers1
2015 Policy-Driven Configuration Management for NoSQL
abstract
NoSQL systems have become the vital components to deliver big data services in the Cloud. However, existing NoSQL systems rely on experienced administrators to configure and tune the wide range of configurable parameters in order to achieve high performance. In this paper, we present a policy-driven configuration management system for NoSQL systems, called PCM. PCM can identify workload sensitive configuration parameters and capture the tuned parameters for different workloads as configuration policies. PCM also can be used to analyze the range of configuration parameters that may impact on the runtime performance of NoSQL systems in terms of read and write workloads. The configuration optimization recommended by PCM can enable NoSQL systems such as HBase to run much more efficiently than the default settings for both individual worker node and entire cluster in the Cloud. Our experimental results show that HBase under the PCM configuration outperforms the default configuration and some simple configurations on a range of workloads with offering significantly higher throughput.
Ling Liu 0001, Nong Xiao 0001, Yang Zhou 0001, Qi Zhang 0009
CLOUD5
2015 Shared Memory Optimization in Virtualized Cloud
abstract
Shared memory management is widely recognized as an optimization technique in the virtualized cloud. Most current shared memory techniques allocate shared memory resources from guest VMs based on pre-defined system configurations. Such static management of shared memory not only increases the VM memory pressure, but also limits the flexibility to balance shared memory resources across multiple VMs running on a single host. In this paper, we present a dynamic shared memory management framework that enables multiple VMs to dynamically access shared memory resources according to their demands. We illustrate our system design through two case studies: One aims to improve the performance of inter-domain communication while the other aims to improve VM memory swapping efficiency. We demonstrate that the dynamic shared memory mechanism not only improves the utilization of shared memory resources but also significantly enhances the performance of VM applications. Our experimental results show that by using dynamic shared memory management, we can improve the performance of inter-VM communication by up to 45 times, while mitigating the VM memory swapping overhead by up to 58%.
Qi Zhang 0009, Ling Liu 0001
CLOUD1
2015 Fast Iterative Graph Computation with Resource Aware Graph Parallel Abstractions
abstract
Iterative computation on large graphs has challenged system research from two aspects: (1) how to conduct high performance parallel processing for both in-memory and out-of-core graphs; and (2) how to handle large graphs that exceed the resource boundary of traditional systems by resource aware graph partitioning such that it is feasible to run large-scale graph analysis on a single PC. This paper presents GraphLego, a resource adaptive graph processing system with multi-level programmable graph parallel abstractions. GraphLego is novel in three aspects: (1) we argue that vertex-centric or edge-centric graph partitioning are ineffective for parallel processing of large graphs and we introduce three alternative graph parallel abstractions to enable a large graph to be partitioned at the granularity of subgraphs by slice, strip and dice based partitioning; (2) we use dice-based data placement algorithm to store a large graph on disk by minimizing non-sequential disk access and enabling more structured in-memory access; and (3) we dynamically determine the right level of graph parallel abstraction to maximize sequential access and minimize random access. GraphLego can run efficiently on different computers with diverse resource capacities and respond to different memory requirements by real-world graphs of different complexity. Extensive experiments show the competitiveness of GraphLego against existing representative graph processing systems, such as GraphChi, GraphLab and X-Stream.
Yang Zhou 0001, Ling Liu 0001, Kisung Lee, Calton Pu, Qi Zhang 0009
HPDC5
2015 Clustering Service Networks with Entity, Attribute, and Link Heterogeneity
abstract
Many popular web service networks are content-rich in terms of heterogeneous types of entities and links, associated with incomplete attributes. Clustering such heterogeneous service networks demands new clustering techniques that can handle two heterogeneity challenges: (1) multiple types of entities co-exist in the same service network with multiple attributes, and (2) links between entities have diverse types and carry different semantics. Existing heterogeneous graph clustering techniques tend to pick initial centroids uniformly at random, specify the number k of clusters in advance, and fix k during the clustering process. In this paper, we propose Service Cluster, a novel heterogeneous service network clustering algorithm with four unique features. First, we incorporate various types of entity, attribute and link information into a unified distance measure. Second, we design a Discrete Steepest Descent method to naturally produce initial k and initial centroids simultaneously. Third, we propose a dynamic learning method to automatically adjust the link weights towards clustering convergence. Fourth, we develop an effective optimization strategy to identify new suitable k and k well-chosen centroids at each clustering iteration. Extensive evaluation on real datasets demonstrates that Service Cluster outperforms existing representative methods in terms of both effectiveness and efficiency.
Yang Zhou 0001, Ling Liu 0001, Calton Pu, Kisung Lee, Balaji Palanisamy, Emre Yigitoglu, Qi Zhang 0009
ICWS8
2015 Scaling iterative graph computations with GraphMap
abstract
In recent years, systems researchers have devoted considerable effort to the study of large-scale graph processing. Existing distributed graph processing systems such as Pregel, based solely on distributed memory for their computations, fail to provide seamless scalability when the graph data and their intermediate computational results no longer fit into the memory; and most distributed approaches for iterative graph computations do not consider utilizing secondary storage a viable solution. This paper presents GraphMap, a distributed iterative graph computation framework that maximizes access locality and speeds up distributed iterative graph computations by effectively utilizing secondary storage. GraphMap has three salient features: (1) It distinguishes data states that are mutable during iterative computations from those that are read-only in all iterations to maximize sequential access and minimize random access. (2) It entails a two-level graph partitioning algorithm that enables balanced workloads and locality-optimized data placement. (3) It contains a proposed suite of locality-based optimizations that improve computational efficiency. Extensive experiments on several real-world graphs show that GraphMap outperforms existing distributed memory-based systems for various iterative graph algorithms.
Kisung Lee, Ling Liu 0001, Karsten Schwan, Calton Pu, Qi Zhang 0009, Yang Zhou 0001, Emre Yigitoglu, Pingpeng Yuan
SC5
2015 GraphTwist: Fast Iterative Graph Computation with Two-tier Optimizations
abstract
Large-scale real-world graphs are known to have highly skewed vertex degree distribution and highly skewed edge weight distribution. Existing vertex-centric iterative graph computation models suffer from a number of serious problems: (1) poor performance of parallel execution due to inherent workload imbalance at vertex level; (2) inefficient CPU resource utilization due to short execution time for low-degree vertices compared to the cost of in-memory or on-disk vertex access; and (3) incapability of pruning insignificant vertices or edges to improve the computational performance. In this paper, we address the above technical challenges by designing and implementing a scalable, efficient, and provably correct two-tier graph parallel processing system, GraphTwist. At storage and access tier, GraphTwist maximizes parallel efficiency by employing three graph parallel abstractions for partitioning a big graph by slice, strip or dice based partitioning techniques. At computation tier, GraphTwist presents two utility-aware pruning strategies: slice pruning and cut pruning, to further improve the computational performance while preserving the computational utility defined by graph applications. Theoretic analysis is provided to quantitatively prove that iterative graph computations powered by utility-aware pruning techniques can achieve a very good approximation with bounds on the introduced error.
Yang Zhou 0001, Ling Liu 0001, Kisung Lee, Qi Zhang 0009
Proc. VLDB Endow.4
2014 Improving Hadoop Service Provisioning in a Geographically Distributed Cloud
abstract
With more data generated and collected in a geographically distributed manner, combined by the increased computational requirements for large scale data-intensive analysis, we have witnessed the growing demand for geographically distributed Cloud datacenters and hybrid Cloud service provisioning, enabling organizations to support instantaneous demand of additional computational resources and to expand inhouse resources to maintain peak service demands by utilizing cloud resources. A key challenge for running applications in such a geographically distributed computing environment is how to efficiently schedule and perform analysis over data that is geographically distributed across multiple datacenters. In this paper, we first compare multi-datacenter Hadoop deployment with single-datacenter Hadoop deployment to identify the performance issues inherent in a geographically distributed cloud. A generalization of the problem characterization in the context of geographically distributed cloud datacenters is also provided with discussions on general optimization strategies. Then we describe the design and implementation of a suite of system-level optimizations for improving performance of Hadoop service provisioning in a geo-distributed cloud, including prediction-based job localization, configurable HDFS data placement, and data prefetching. Our experimental evaluation shows that our prediction based localization has very low error ratio, smaller than 5%, and our optimization can improve the execution time of Reduce phase by 48.6%.
Qi Zhang 0009, Ling Liu 0001, Kisung Lee, Yang Zhou 0001, Aameek Singh, NagaPramod Mandagere, Sandeep Gopisetty, Gabriel Alatorre
IEEE CLOUD1
2014 Improving MapReduce Performance in a Heterogeneous Cloud: A Measurement Study
abstract
Hybrid clouds, geo-distributed cloud and continuous upgrades of computing, storage and networking resources in the cloud have driven datacenters evolving towards heterogeneous clusters. Unfortunately, most of MapReduce implementations are designed for homogeneous computing environments and perform poorly in heterogeneous clusters. Although a fair of research efforts have dedicated to improve MapReduce performance, there still lacks of in-depth understanding of the key factors that affect the performance of MapReduce jobs in heterogeneous clusters. In this paper, we present an extensive experimental study on two categories of factors: system configuration and task scheduling. Our measurement study shows that an in-depth understanding of these factors is critical for improving MapReduce performance in a heterogeneous environment. We conclude with five key findings: (1) Early shuffle, though effective for reducing the latency of MapReduce jobs, can impact the performance of map tasks and reduce tasks differently when running on different types of nodes. (2) Two phases in map tasks have different sensitive to input block size and the ratio of sort phase with different block size is different for different type of nodes. (3) Scheduling map or reduce tasks dynamically with node capacity and workload awareness can further enhance the job performance and improve resource consumption efficiency. (4) Although random scheduling of reduce tasks works well in homogeneous clusters, it can significantly degrade the performance in heterogeneous clusters when shuffled data size is large. (5) Phase-aware progress rate estimation and speculation strategy can provide substantial performance gain over the state of art speculation scheduler.
Ling Liu 0001, Qi Zhang 0009, Xiaoshe Dong
IEEE CLOUD3
2014 HConfig: Resource adaptive fast bulk loading in HBase
abstract
NoSQL (Not only SQL) data stores become a vital component in many big data computing platforms due to its inherent horizontal scalability. HBase is an open-source distributed NoSQL store that is widely used by many Internet enterprises to handle their big data computing applications (e.g. Facebook h
Ling Liu 0001, Nong Xiao 0001, Fang Liu 0002, Qi Zhang 0009
CollaborateCom5
2014 e-PPI: Locator Service in Information Networks with Personalized Privacy Preservation
abstract
In emerging information networks, having a privacy preserving index (or PPI) is critically important for locating information of interest for data sharing across autonomous providers while preserving privacy. An understudied problem for PPI techniques is how to provide controllable privacy preservation, given the innate difference of privacy concerns regarding different data owners. In this paper we present a personalized privacy preserving index, coined ε-PPI, which guarantees quantitative privacy preservation differentiated by personal identities. We devise a new common-identity attack that breaks existing PPI's and propose an identity-mixing protocol against the attack in ε-PPI. The proposed ε-PPI construction protocol is the first without any trusted third party and/or trust relationships between providers. We have implemented our ε-PPI construction protocol by using generic MPC techniques (secure multi-party computation) and optimized the performance to a practical level by minimizing the expensive MPC part.
Yuzhe Tang, Ling Liu 0001, Arun Iyengar, Kisung Lee, Qi Zhang 0009
ICDCS5
2013 Efficient and Customizable Data Partitioning Framework for Distributed Big RDF Data Processing in the Cloud
abstract
Big data business can leverage and benefit from the Clouds, the most optimized, shared, automated, and virtualized computing infrastructures. One of the important challenges in processing big data in the Clouds is how to effectively partition the big data to ensure efficient distributed processing of the data. In this paper we present a Scalable and yet customizable data PArtitioning framework, called SPA, for distributed processing of big RDF graph data. We choose big RDF datasets as our focus of the investigation for two reasons. First, the Linking Open Data cloud has put forwards a good number of big RDF datasets with tens of billions of triples and hundreds of millions of links. Second, such huge RDF graphs can easily overwhelm any single server due to the limited memory and CPU capacity and exceed the processing capacity of many conventional data processing software systems. Our data partitioning framework has two unique features. First, we introduce a suite of vertexcentric data partitioning building blocks to allow efficient and yet customizable partitioning of large heterogeneous RDF graph data. By efficient, we mean that the SPA data partitions can support fast processing of big data of different sizes and complexity. By customizable, we mean that the SPA partitions are adaptive to different query types. Second, we propose a selection of scalable techniques to distribute the building block partitions across a cluster of compute nodes in a manner that minimizes inter-node communication cost by localizing most of the queries on distributed partitions. We evaluate our data partitioning framework and algorithms through extensive experiments using both benchmark and real datasets. Our experimental results show that the SPA data partitioning framework is not only efficient for partitioning and distributing big RDF datasets of diverse sizes and structures but also effective for processing big data queries of different types and complexity.
Kisung Lee, Ling Liu 0001, Yuzhe Tang, Qi Zhang 0009, Yang Zhou 0001
IEEE CLOUD4
2013 Residency-Aware Virtual Machine Communication Optimization: Design Choices and Techniques
abstract
Network I/O workloads are dominating in many data centers and cloud computing environments today. One way to improve inter Virtual Machine (VM) communication efficiency is to support co-resident VM communication by using shared memory based approaches and to resort to the traditional TCP/IP for inter-VM communications between VMs that are located on different physical hosts. Although a number of independent efforts are dedicated to improving communication efficiency between co-resident VMs, they differ from one another in terms of how the inter-VM communication optimization is carried out and where in the software stack the shared memory channel is established. In this paper, we provide an in-depth overview of the design choices and techniques for optimizing the performance of the co-resident inter-VM communication, with dual objectives. First, we describe the core design guidelines and key issues for optimizing inter-VM communication by using shared memory based mechanisms. Typical issues include choices of implementation layer in the software stack, seamless agility for VM live migration and VM dynamic deployment support, multilevel transparency. Second, we conduct a comprehensive analysis of representative state-of-the-art research efforts and implementation techniques based on the core design guidelines. We also give an analysis of future requirements in advanced features such as reliability, security and stability. The research reported in this paper not only provides the reference for developing the next generation of inter-VM communication optimization mechanisms, but also offers opportunities for both cloud infrastructure providers and cloud service consumers to improve inter-VM communication efficiency in virtualized platforms.
Yi Ren 0008, Ling Liu 0001, Qi Zhang 0009, Qingbo Wu 0003, Jinzhu Kong, Jianbo Guan, Huadong Dai
IEEE CLOUD3
2013 Residency Aware Inter-VM Communication in Virtualized Cloud: Performance Measurement and Analysis
abstract
A known problem for virtualized cloud data centers is the inter-VM communication inefficiency for data transfer between co-resident VMs. Several engineering efforts have been made on building a shared memory based channel between co-resident VMs. The implementations differ in terms of whether user/program transparency, OS kernel transparency or VMM transparency is supported. However, none of existing works has engaged in an in-depth measurement study with quantitative and qualitative analysis on performance improvements as well as tradeoffs introduced by such a residency-aware inter-VM communication mechanism. In this paper we present an extensive experimental study, aiming at addressing a number of fundamental issues and providing deeper insights regarding the design of a shared memory channel for co-resident VMs. Example questions include how much performance gains can a residency-aware shared memory inter-VM communication mechanism provide under different mixtures of local and remote network I/O workloads, what overhead will the residence-awareness detection and communication channel switch introduce over the remote inter-VM communication, what factors may exert significant impact on the throughput and latency performance of such a shared memory channel. We believe that this measurement study not only helps system developers to gain valuable lessons and generate new ideas to further improve the inter-VM communication performance. It also offers new opportunities for cloud service providers to deploy their services more efficiently and for cloud service consumers to improve the performance of their application systems running in the Cloud.
Qi Zhang 0009, Ling Liu 0001, Yi Ren 0008, Kisung Lee, Yuzhe Tang, Yang Zhou 0001
IEEE CLOUD1
2013 A New Disk I/O Model of Virtualized Cloud Environment
abstract
In a traditional virtualized cloud environment, using asynchronous I/O in the guest file system and synchronous I/O in the host file system to handle an asynchronous user disk write exhibits several drawbacks, such as performance disturbance among different guests and consistency maintenance across guest failures. To improve these issues, this paper introduces a novel disk I/O model for virtualized cloud system called HypeGear, where the guest file system uses synchronous operations to deal with the guest write request and the host file system performs asynchronous operations to write the data to the hard disk. A prototype system is implemented on the Xen hypervisor and our experimental results verify that this new model has many advantages over the conventional asynchronous-synchronous model. We also evaluate the overhead of asynchronous I/O at host, which is brought by our new model. The result demonstrates that it enforces little cost on host layer.
Dingding Li, Xiaofei Liao, Hai Jin 0001, Bing Bing Zhou, Qi Zhang 0009
IEEE Trans. Parallel Distributed Syst.5