EDBT 2026 Demo / reviewers in the wild / expert
Tian Guo 0001
dblp:55/3523-1
· DBLP profile ↗
50ranked-venue papers
9as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 5 since 2021Computer networks · 5 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ARBot: A High-Fidelity Robotic Manipulator Teleoperation Framework for Human-Centered Augmented Reality EvaluationabstractValidating Augmented Reality (AR) tracking and interaction models requires precise, repeatable ground-truth motion. However, human users cannot reliably perform consistent motion due to biome-chanical variability. Robotic manipulators are promising to act as human motion proxies if they can mimic human movements. In this work, we design and implement ARBot, a real-time teleoperation platform that can effectively capture natural human motion and accurately replay the movements via robotic manipulators. ARBot includes two capture models: stable wrist motion capture via a custom CV and IMU pipeline, and natural 6-DOF control via a mobile application. We design a proactively-safe QP controller to ensure smooth, jitter-free execution of the robotic manipulator, enabling it to function as a high-fidelity record and replay physical proxy. We open-source ARBot and release a benchmark dataset of 132 human and synthetic trajectories captured using ARBot to support controllable and scalable AR evaluation. Harsh Chhajed, Tian Guo 0001 |
MMSys | 2 |
| 2026 | AR as an Evaluation Playground: Bridging Metric and Visual Perception of Computer Vision Models
Ashkan Ganj, Yiqin Zhao, Tian Guo 0001 |
MMSys | 3 |
| 2025 | CarbonDIS: Carbon-Aware DNN Inference Scheduling on Heterogeneous GPUsabstractDeploying deep neural network (DNN) inference applications in data centers and clouds to empower various services is on the increase. The significant power consumption of these applications contributes to the carbon emissions of underlying infrastructures. Therefore, minimizing the carbon emissions of these applications leads to a reduced carbon footprint of data centers and clouds. This paper introduces CarbonDIS, a carbon-aware inference scheduler designed to maintain the latency of DNN inference applications while minimizing their carbon emissions by leveraging heterogeneous GPUs. CarbonDIS considers an inference architecture where high-end and low-end GPUs are used in tandem to serve inference jobs. By leveraging the varying computing capability and power consumption of heterogeneous GPUs, CarbonDIS can effectively balance the performance and carbon emissions. Evaluation using three types of GPUs, 10 DNN models, and real-world carbon intensity traces show that CarbonDIS can find a trade-off between carbon footprint and latency of the jobs and reduce the carbon emissions by 16% compared with a performance-centric approach. Seyed Morteza Nabavinejad, Shubbhi Taneja, Tian Guo 0001 |
CCGrid | 4 |
| 2025 | Carbon-Efficient Internet Video StreamingabstractCarbon-efficient computing has drawn significant attention recently aiming to achieve environmental sustainability. Several computation intensive applications, such as deep learning, have been studied in the community for carbon optimizations. However, video streaming, one of the most popular and resource-consuming Internet applications, is under-explored for carbon efficiency. In this paper, we for the first time investigate the carbon efficiency of video streaming systems with the goal of reducing carbon emissions caused by hosting the video streaming service. We develop a dynamic workload migration mechanism utilizing real-time carbon intensity data to select the hosting data center. Furthermore, to minimize the impact on the end-user experience, we consider the migration frequency to maintain service stability when making the migration decisions. The evaluation results indicate significant carbon reductions and acceptable stream switching overhead. Tian Guo 0001, Sheng Wei 0001 |
MMSP | 2 |
| 2025 | HybridDepth: Robust Metric Depth Fusion by Leveraging Depth from Focus and Single-Image PriorsabstractWe propose Hybriddepth, a robust depth estimation pipeline that addresses key challenges in depth estimation, including scale ambiguity, hardware heterogene-ity, and generalizability. Hybriddepth leverages focal stack, data conveniently accessible in common mobile de-vices, to produce accurate metric depth maps. By incorpo-rating depth priors afforded by recent advances in single-image depth estimation, our model achieves a higher level of structural detail compared to existing methods. We test our pipeline as an end-to-end system, with a newly developed mobile client to capture focal stacks, which are then sent to a GPU-powered server for depth estimation. Comprehensive quantitative and qualitative analyses demonstrate that Hybriddepth outperforms state-of-the-art (SOTA) models on common datasets such as DDFF12 and NYU Depth V2. Hybriddepth also shows strong zero-shot generalization. When trained on NYU Depth V2, Hybriddepth surpasses SOTA models in zero-shot per-formance on ARKitScenes and delivers more structurally accurate depth maps on Mobile Depth. The code is avail-able at https://github.com/cake-labIHybridDepth/. Ashkan Ganj, Hang Su 0005, Tian Guo 0001 |
WACV | 3 |
| 2025 | BridgeGC: An Efficient Cross-Level Garbage Collector for Big Data FrameworksabstractPopular big data frameworks commonly run atop Java Virtual Machine (JVM) and rely on garbage collection (GC) mechanism to automatically allocate/reclaim in-memory objects. Existing garbage collectors are designed based on the hypothesis that most objects are short lived. However, big data frameworks usually generate many long-lived data objects, which can cause heavy GC overhead. Recent approaches have reduced GC overhead in big data frameworks but still suffer from heavy human efforts, additional runtime overhead, or suboptimal GC efficiency. This article describes the design of BridgeGC , a big-data-friendly garbage collector that significantly reduces GC overhead introduced by long-lived data objects. BridgeGC follows a cross-level co-design. At the big data framework level, BridgeGC provides two annotations for framework developers to denote the creation and release of data objects. Based on the annotations, BridgeGC tracks the lifecycles of annotated data objects and optimizes their allocation/reclamation at the GC level. At the GC level, we design a label-based allocator that stores data objects separately from other objects and balances their memory usage in the same JVM, leading to fewer GC cycles. We further design an efficient collector to eliminate unnecessary marking and copying of data objects during GC cycles, lowering the GC time. We have integrated BridgeGC into OpenJDK ZGC. The extensive evaluation, using two popular big data frameworks (Flink and Spark) and a key–value database (Cassandra), shows that BridgeGC achieves 31–82% GC time reduction compared to the baseline ZGC. BridgeGC also outperforms other traditional and academic garbage collectors in end-to-end performance. Lijie Xu, Tian Guo 0001, Wensheng Dou, Hongbin Zeng, Wei Wang 0049, Jun Wei 0001, Tao Huang 0001 |
ACM Trans. Archit. Code Optim. | 3 |
| 2024 | MediatorDNN: Contention Mitigation for Co-Located DNN Inference JobsabstractWith the increase in computing power of cutting-edge hardware platforms, it is a common practice to run multiple jobs on a single machine for improved resource utilization and throughput. However, this leads to inevitable resource contention among co-located jobs, impacting their performance. The resource contention can worsen due to fluctuations in resource utilization of jobs caused by variations in their input workload. To tackle the co-location contention for DNN inference jobs, we propose MediatorDNN, which considers contention and resource utilization variation when co-locating DNN inference jobs. It profiles each DNN, monitors microarchitectural metrics such as memory bandwidth and cache access pattern, and high-level resource utilization like CPU utilization. Based on profiling results and leveraging Modern Portfolio Theory (MPT), MediatorDNN decides on the co-location of jobs. Experimental results with various DNNs on two hardware platforms show that MediatorDNN improves throughput by up to 108% (21 % on average) compared to an approach only considering contention and ignoring resource utilization variation. Seyed Morteza Nabavinejad, Sherief Reda, Tian Guo 0001 |
CLOUD | 3 |
| 2024 | Towards In-context Environment Sensing for Mobile Augmented RealityabstractEnvironment sensing is a fundamental task in mobile augmented reality (AR). However, on-device sensing and computing resources often limit mobile AR sensing capability, making high-quality environment sensing challenging to achieve. In recent years, in-context sensing, a new sensing system design paradigm, has emerged with the promise of achieving accurate, efficient, and robust sensing results. In this work, we first formally define the in-context sensing design paradigm. We summarize its primary challenges as the uncertainty of environmental information availability. To quantify the impact of sensing context data, we present two in-depth case studies that show how it can impact different aspects of mobile AR sensing systems. Yiqin Zhao, Ashkan Ganj, Tian Guo 0001 |
MobiCom | 3 |
| 2024 | CE-NAS: An End-to-End Carbon-Efficient Neural Architecture Search FrameworkabstractThis work presents a novel approach to neural architecture search (NAS) that aims to increase carbon efficiency for the model design process. The proposed framework CE-NAS addresses the key challenge of high carbon cost associated with NAS by exploring the carbon emission variations of energy and energy differences of different NAS algorithms. At the high level, CE-NAS leverages a reinforcement-learning agent to dynamically adjust GPU resources based on carbon intensity, predicted by a time-series transformer, to balance energy-efficient sampling and energy-intensive evaluation tasks. Furthermore, CE-NAS leverages a recently proposed multi-objective optimizer to effectively reduce the NAS search space. We demonstrate the efficacy of CE-NAS in lowering carbon emissions while achieving SOTA results for both NAS datasets and open-domain NAS tasks. For example, on the HW-NasBench dataset, CE-NAS reduces carbon emissions by up to 7.22X while maintaining a search efficiency comparable to vanilla NAS. For open-domain NAS tasks, CE-NAS achieves SOTA results with 97.35% top-1 accuracy on CIFAR-10 with only 1.68M parameters and a carbon consumption of 38.53 lbs of CO2. On ImageNet, our searched model achieves 80.6% top-1 accuracy with a 0.78 ms TensorRT latency using FP16 on NVIDIA V100, consuming only 909.86 lbs of CO2, making it comparable to other one-shot-based NAS baselines. Our code is available at https://github.com/cake-lab/CE-NAS. Yunzhuo Liu, Bo Jiang 0003, Tian Guo 0001 |
NeurIPS | 4 |
| 2024 | Multi-Objective Neural Architecture Search by Learning Search Space PartitionsabstractDeploying deep learning models requires taking into consideration neural network metrics such as model size, inference latency, and #FLOPs, aside from inference accuracy. This results in deep learning model designers leveraging multi-objective optimization to design effective deep neural networks in multiple criteria. However, applying multi-objective optimizations to neural architecture search (NAS) is nontrivial because NAS tasks usually have a huge search space, along with a non-negligible searching cost. This requires effective multi-objective search algorithms to alleviate the GPU costs. In this work, we implement a novel multi-objectives optimizer based on a recently proposed meta-algorithm called LaMOO on NAS tasks. In a nutshell, LaMOO speedups the search process by learning a model from observed samples to partition the search space and then focusing on promising regions likely to contain a subset of the Pareto frontier. Using LaMOO, we observe an improvement of more than 200% sample efficiency compared to Bayesian optimization and evolutionary-based multi-objective optimizers on different NAS datasets. For example, when combined with LaMOO, qEHVI achieves a 225% improvement in sample efficiency compared to using qEHVI alone in NasBench201. For real-world tasks, LaMOO achieves 97.36% accuracy with only 1.62M #Params on CIFAR10 in only 600 search samples. On ImageNet, our large model reaches 80.4% top-1 accuracy with only 522M #FLOPs. Linnan Wang, Tian Guo 0001 |
J. Mach. Learn. Res. | 3 |
| 2023 | Measuring the Impact of Gradient Accumulation on Cloud-based Distributed TrainingabstractGradient accumulation (GA) is a commonly adopted technique for addressing the GPU memory shortage problem in model training. It reduces memory consumption at the cost of increased computation time. Although widely used, its benefits to model training have not been systematically studied. Our work evaluates and summarizes the benefits of GA, especially in cloud-based distributed training scenarios, where training cost is determined by both execution time and resource consumption. We focus on how GA can be utilized to balance execution time and resource consumption to achieve the lowest bills. Through empirical evaluations on AliCloud platforms, we observe that the total training cost can be reduced by 31.2% on average with a 17.3% increase in training time, when GA is introduced in the large-model and small-bandwidth scenarios with data-parallel training strategies. Besides, taking micro-batch size into optimization can further decrease training time and cost by 21.2% and 24.8% on average, respectively, for hybrid-parallel strategies in large-model and GPU training scenarios. Zimeng Huang, Bo Jiang 0003, Tian Guo 0001, Yunzhuo Liu |
CCGrid | 3 |
| 2023 | Layercake: Efficient Inference Serving with Cloud and Mobile ResourcesabstractMany mobile applications are now integrating deep learning models into their core functionality. These functionalities have diverse latency requirements while demanding high-accuracy results. Currently, mobile applications statically decide to use either in-cloud inference, relying on a fast and consistent network, or on-device execution, relying on sufficient local resources. However, neither mobile networks nor computation resources deliver consistent performance in practice. Consequently, mobile inference often experiences variable performance or struggles to meet performance goals, when inference execution decisions are not made dynamically. In this paper, we introduce Layer Cake, a deep-learning inference framework that dynamically selects the best model and location for executing inferences. Layercake accomplishes this by tracking model state and availability, both locally and remotely, as well as the network bandwidth, allowing for accurate estimations of model response time. By doing so, Layercake achieves latency targets in up to 96.4% of cases, which is an improvement of 16.7% over similar systems, while decreasing the cost of cloud-based resources by over 68.33% than in-cloud inference. Samuel S. Ogden, Tian Guo 0001 |
CCGrid | 2 |
| 2022 | Multi-objective Optimization by Learning Space Partition
Linnan Wang, Kevin Yang, Tianjun Zhang, Tian Guo 0001, Yuandong Tian |
ICLR | 5 |
| 2022 | Privacy-preserving Reflection Rendering for Augmented RealityabstractWhen the virtual objects consist of reflective materials, the required lighting information to render such objects can consist of privacy-sensitive information outside the current camera view. In this paper, we show, for the first time, that accuracy-driven multi-view environment lighting can reveal out-of-camera scene information and compromise privacy. We present a simple yet effective privacy attack that extracts sensitive scene information such as human faces and text from rendered objects under several application scenarios. Yiqin Zhao, Sheng Wei 0001, Tian Guo 0001 |
ACM Multimedia | 3 |
| 2022 | Power-efficient live virtual reality streaming using edge offloadingabstractThis paper aims to address the significant power challenges in live virtual reality (VR) streaming (a.k.a., 360-degree video streaming), where the VR view rendering and the advanced deep learning operations (e.g., super-resolution) consume a considerable amount of power draining the battery-constrained VR headset. We develop EdgeVR, a power optimization technique for live VR streaming, which offloads the on-device VR rendering and deep learning operations to an edge server for power savings. To address the significantly increased motion-to-photon (MtoP) latency due to the edge offloading, we develop a live VR viewport prediction method to pre-render the VR views on the edge server and compensate for the round-trip delays. We evaluate the effectiveness of EdgeVR using an end-to-end live VR streaming system with an empirical VR head movement dataset involving 48 users watching 9 VR videos. The results reveal that EdgeVR achieves power-efficient live VR streaming with low MtoP latency. Xianglong Feng, Zhongze Tang, Nan Jiang 0020, Tian Guo 0001, Lisong Xu, Sheng Wei 0001 |
NOSSDAV | 5 |
| 2021 | FiShNet: Fine-Grained Filter Sharing for Resource-Efficient Multi-Task LearningabstractMulti-task learning has attracted much attention in recent years, where the goal is to learn multiple tasks by exploiting the similarities and differences between the tasks. Previous researches on multi-task learning mainly focus on flexible methods for feature sharing (e.g., soft sharing) under resource-sufficient settings (e.g., on GPU servers). However, in many real-world applications, we often need to deploy multi-task learning models on resource-constrained platforms (e.g., mobile devices). The high resource requirement of soft-sharing methods can make them hard to deploy on mobile devices. In this paper, we study the problem of Resource-efficient Multi-Task Learning (MTL), where the goal is to design a resource-friendly model that suits resource-constrained inference environment, e.g., security camera or mobile devices. We formulate the Resource-efficient MTL problem as a fine-grained filter sharing problem, i.e., learning how to share filters at any given convolutional layers among multiple tasks. We proposed a novel solution for parameter sharing, called FiShNet. Different from soft-sharing approaches, where the computational cost per task is growing w.r.t. the number of other tasks, FiShNet can achieve high accuracy comparable to soft-sharing approaches, while only consuming a constant computational cost per task. Different from hard-sharing approaches, where the parameter sharing structures are hand-picked, FiShNet can learn how to share parameters directly on the training data with finer-grained sharing. We evaluate FiShNet on a number of problem settings and datasets for multi-task learning. We show that FiShNet achieves high accuracy when compared with state-of-the-art methods in multi-task learning, while only requiring a fraction of the computational resource. Xiangnan Kong, Tian Guo 0001, Xinlu He |
CIKM | 3 |
| 2021 | Enabling Sustainable Clouds: The Case for Virtualizing the Energy SystemabstractCloud platforms' growing energy demand and carbon emissions are raising concern about their environmental sustainability. The current approach to enabling sustainable clouds focuses on improving energy-efficiency and purchasing carbon offsets. These approaches have limits: many cloud data centers already operate near peak efficiency, and carbon offsets cannot scale to near zero carbon where there is little carbon left to offset. Instead, enabling sustainable clouds will require applications to adapt to when and where unreliable low-carbon energy is available. Applications cannot do this today because their energy use and carbon emissions are not visible to them, as the energy system provides the rigid abstraction of a continuous, reliable energy supply. This vision paper instead advocates for a "carbon first" approach to cloud design that elevates carbon-efficiency to a firs--class metric. To do so, we argue that cloud platforms should virtualize the energy system by exposing visibility into, and software-defined control of, it to applications, enabling them to define their own abstractions for managing energy and carbon emissions based on their own requirements. Noman Bashir, Tian Guo 0001, Mohammad Hajiesmaili, David Irwin 0001, Prashant J. Shenoy, Ramesh K. Sitaraman, Abel Souza, Adam Wierman |
SoCC | 2 |
| 2021 | On the Future of Cloud EngineeringabstractEver since the commercial offerings of the Cloud started appearing in 2006, the landscape of cloud computing has been undergoing remarkable changes with the emergence of many different types of service offerings, developer productivity enhancement tools, and new application classes as well as the manifestation of cloud functionality closer to the user at the edge. The notion of utility computing, however, has remained constant throughout its evolution, which means that cloud users always seek to save costs of leasing cloud resources while maximizing their use. On the other hand, cloud providers try to maximize their profits while assuring service-level objectives of the cloud-hosted applications and keeping operational costs low. All these outcomes require systematic and sound cloud engineering principles. The aim of this paper is to highlight the importance of cloud engineering, survey the landscape of best practices in cloud engineering and its evolution, discuss many of the existing cloud engineering advances, and identify both the inherent technical challenges and research opportunities for the future of cloud computing in general and cloud engineering in particular. David Bermbach, Abhishek Chandra, Chandra Krintz, Aniruddha S. Gokhale, Aleksander Slominski, Lauritz Thamsen, Everton Cavalcante, Tian Guo 0001, Ivona Brandic, Richard Wolski |
IC2E | 8 |
| 2021 | Quantifying and Improving Performance of Distributed Deep Learning with Cloud StorageabstractCloud computing provides a powerful yet low-cost environment for distributed deep learning workloads. However, training complex deep learning models often requires accessing large amounts of data, which can easily exceed the capacity of local disks. Prior research often overlooks this training data problem by implicitly assuming that data is available locally or via low latency network-based data storage. Such implicit assumptions often do not hold in a cloud-based training environment, where deep learning practitioners create and tear down dedicated GPU clusters on demand, or do not have the luxury of local storage, such as in serverless workloads. In this work, we investigate the performance of distributed training that leverages training data residing entirely inside cloud storage buckets. These buckets promise low storage costs, but come with inherent bandwidth limitations that make them seem unsuitable for an efficient training solution. To account for these bandwidth limitations, we propose the use of two classical techniques, namely caching and pre-fetching, to mitigate the training performance degradation. We implement a prototype, DELI, based on the popular deep learning framework PyTorch by building on its data loading abstractions. We then evaluate the training performance of two deep learning workloads using Google Cloud's NVIDIA K80 GPU servers and show that we can reduce the time that the training loop is waiting for data by 85.6%-93.5% compared to loading directly from a storage bucket - thus achieving comparable performance to loading data directly from disk - while only storing a fraction of the data locally at a time. In addition, DELI has the potential of lowering the cost of running a training workload, especially on models with long per-epoch training times. Nicholas Krichevsky, Renee St Louis, Tian Guo 0001 |
IC2E | 3 |
| 2021 | Memory-Efficient Deep Learning Inference in Trusted Execution EnvironmentsabstractThis study identifies and proposes techniques to alleviate two key bottlenecks to executing deep neural networks in trusted execution environments (TEEs): page thrashing during the execution of convolutional layers and the decryption of large weight matrices in fully-connected layers. For the former, we propose a novel partitioning scheme, y-plane partitioning, designed to (i) provide consistent execution time when the layer output is large compared to the TEE secure memory; and (ii) significantly reduce the memory footprint of convolutional layers. For the latter, we leverage quantization and compression. In our evaluation, the proposed optimizations incurred latency overheads ranging from 1.09X to 2X baseline for a wide range of TEE sizes; in contrast, an unmodified implementation incurred latencies of up to 26X when running inside of the TEE. Jean-Baptiste Truong, William Gallagher, Tian Guo 0001, Robert J. Walls |
IC2E | 3 |
| 2021 | Sync-Switch: Hybrid Parameter Synchronization for Distributed Deep LearningabstractStochastic Gradient Descent (SGD) has become the de facto way to train deep neural networks in distributed clusters. A critical factor in determining the training throughput and model accuracy is the choice of the parameter synchronization protocol. For example, while Bulk Synchronous Parallel (BSP) often achieves better converged accuracy, the corresponding training throughput can be negatively impacted by stragglers. In contrast, Asynchronous Parallel (ASP) can have higher throughput, but its convergence and accuracy can be impacted by stale gradients. To improve the performance of synchronization protocol, recent work often focuses on designing new protocols with a heavy reliance on hard-to-tune hyper-parameters. In this paper, we design a hybrid synchronization approach that exploits the benefits of both BSP and ASP, i.e., reducing training time while simultaneously maintaining the converged accuracy. Based on extensive empirical profiling, we devise a collection of adaptive policies that determine how and when to switch between synchronization protocols. Our policies include both offline ones that target recurring jobs and online ones for handling transient stragglers. We implement the proposed policies in a prototype system, called Sync-Switch, on top of TensorFlow, and evaluate the training performance with popular deep learning models and datasets. Our experiments show that Sync-Switch can achieve ASP level training speedup while maintaining similar converged accuracy when comparing to BSP. Moreover, Sync-Switch's elastic-based policy can adequately mitigate the impact from transient stragglers. Shijian Li, Oren Mangoubi, Lijie Xu, Tian Guo 0001 |
ICDCS | 4 |
| 2021 | Few-Shot Neural Architecture SearchabstractEfficient evaluation of a network architecture drawn from a large search space remains a key challenge in Neural Architecture Search (NAS). Vanilla NAS evaluates each architecture by training from scratch, which gives the true performance but is extremely time-consuming. Recently, one-shot NAS substantially reduces the computation cost by training only one supernetwork, a.k.a. supernet, to approximate the performance of every architecture in the search space via weight-sharing. However, the performance estimation can be very inaccurate due to the co-adaption among operations. In this paper, we propose few-shot NAS that uses multiple supernetworks, called sub-supernet, each covering different regions of the search space to alleviate the undesired co-adaption. Compared to one-shot NAS, few-shot NAS improves the accuracy of architecture evaluation with a small increase of evaluation cost. With only up to 7 sub-supernets, few-shot NAS establishes new SoTAs: on ImageNet, it finds models that reach 80.5% top-1 accuracy at 600 MB FLOPS and 77.5% top-1 accuracy at 238 MFLOPS; on CIFAR10, it reaches 98.72% top-1 accuracy without using extra data or transfer learning. In Auto-GAN, few-shot NAS outperforms the previously published results by up to 20%. Extensive experiments show that few-shot NAS significantly improves various one-shot methods, including 4 gradient-based and 6 search-based methods on 3 different tasks in NasBench-201 and NasBench1-shot-1. Linnan Wang, Yuandong Tian, Rodrigo Fonseca, Tian Guo 0001 |
ICML | 5 |
| 2021 | Xihe: a 3D vision-based lighting estimation framework for mobile augmented realityabstractOmnidirectional lighting provides the foundation for achieving spatially-variant photorealistic 3D rendering, a desirable property for mobile augmented reality applications. However, in practice, estimating omnidirectional lighting can be challenging due to limitations such as partial panoramas of the rendering positions, and the inherent environment lighting and mobile user dynamics. A new opportunity arises recently with the advancements in mobile 3D vision, including built-in high-accuracy depth sensors and deep learning-powered algorithms, which provide the means to better sense and understand the physical surroundings. Centering the key idea of 3D vision, in this work, we design an edge-assisted framework called Xihe to provide mobile AR applications the ability to obtain accurate omnidirectional lighting estimation in real time. Yiqin Zhao, Tian Guo 0001 |
MobiSys | 2 |
| 2021 | CiNet: Redesigning Deep Neural Networks for Efficient Mobile-Cloud Collaborative InferenceabstractDeep neural networks are increasingly used in end devices such as mobile phones to support novel features, e.g., image classification.Traditional paradigms to support mobile deep inference fall into either cloud-based or on-device-both require access to an entire pre-trained model.As such, the efficacy of mobile deep inference is limited by mobile network conditions and computational capacity.Collaborative inference, a means to splitting inference computation between mobile devices and cloud servers, was proposed to address the limitations of traditional inference through techniques such as image compression or model partition.In this paper, we improve the performance of collaborative inference from a complementary direction, i.e., through redesigning deep neural networks to satisfy the collaboration requirement from the outset.Specifically, we describe the design of a collaboration-aware convolutional neural network, referred to as CiNet, for image classification.CiNet consists of a mobile-side extractor submodel that outputs a small yet relevant patch of the image and a cloud-based submodel that classifies on the image patch.We evaluated the efficiency of CiNet in terms of inference accuracy, computational cost and mobile data transmission on three datasets.Our results demonstrate that CiNet achieved comparable inference accuracy while incurring orders of magnitude less computational cost and 99% less transmitted data, when comparing to both traditional and collaborative inference approaches. Xiangnan Kong, Tian Guo 0001, Yixian Huang |
SDM | 3 |
| 2021 | PieSlicer: Dynamically Improving Response Time for Cloud-based CNN InferenceabstractExecuting deep-learning inference on cloud servers enables the usage of high complexity models for mobile devices with limited resources. However, pre-execution time-the time it takes to prepare and transfer data to the cloud-is variable and can take orders of magnitude longer to complete than inference execution itself. This pre-execution time can be reduced by dynamically deciding the order of two essential steps, preprocessing and data transfer, to better take advantage of on-device resources and network conditions. In this work, we present PieSlicer, a system for making dynamic preprocessing decisions to improve cloud inference performance using linear regression models. PieSlicer then leverages these models to select the appropriate preprocessing location. We show that for image classification applications PieSlicer reduces median and 99th percentile pre-execution time by up to 50.2ms and 217.2ms respectively when compared to static preprocessing methods. Samuel S. Ogden, Xiangnan Kong, Tian Guo 0001 |
ICPE | 3 |
| 2020 | EPNet: Learning to Exit with Flexible Multi-Branch NetworkabstractDynamic inference is an emerging technique that reduces the computational cost of deep neural network under resource-constrained scenarios, such as inference on mobile devices. One way to achieve dynamic inference is to leverage multi-branch neural networks that apply different computation on input data by following different branches. Conventional research on multi-branch neural networks mainly targeted at improving the accuracy of each branch, and use manually designed rules to decide which input follows which branch of the network. Furthermore, these networks often provide a small number of exits, limiting their ability to adapt to external changes. In this paper, we investigate the problem of designing a flexible multi-branch network and early-exiting policies that can adapt to the resource consumption to individual inference request without impacting the inference accuracy. We propose a lightweight branch structure that also provides fine-grained flexibility for early-exiting and leverage Markov decision process (MDP) to automatically learn the early-exiting policies. Our proposed model, EPNet, was effective in reducing inference cost without impacting accuracy by choosing the most suitable branch exit. We also observe that EPNet achieved 3% higher accuracy with an inference budget, compared to state-of-the-art approaches. Xiangnan Kong, Tian Guo 0001 |
CIKM | 3 |
| 2020 | PointAR: Efficient Lighting Estimation for Mobile Augmented Reality
Yiqin Zhao, Tian Guo 0001 |
ECCV (23) | 2 |
| 2020 | PERSEUS: Characterizing Performance and Cost of Multi-Tenant Serving for CNN ModelsabstractDeep learning models are increasingly used for end-user applications, supporting both novel features such as facial recognition, and traditional features, e.g. web search. To accommodate high inference throughput, it is common to host a single pre-trained Convolutional Neural Network (CNN) in dedicated cloud-based servers with hardware accelerators such as Graphics Processing Units (GPUs). However, GPUs can be orders of magnitude more expensive than traditional Central Processing Unit (CPU) servers. These resources could also be under-utilized facing dynamic workloads, which may result in inflated serving costs. One potential way to alleviate this problem is by allowing hosted models to share the underlying resources, which we refer to as multi-tenant inference serving. One of the key challenges is maximizing the resource efficiency for multi-tenant serving given hardware with diverse characteristics, models with unique response time Service Level Agreement (SLA), and dynamic inference workloads. In this paper, we present PERSEUS, a measurement framework that provides the basis for understanding the performance and cost trade-offs of multi-tenant model serving. We implemented PERSEUS in Python atop a popular cloud inference server called Nvidia TensorRT Inference Server. Leveraging PERSEUS, we evaluated the inference throughput and cost for serving various models and demonstrated that multi-tenant model serving led to up to 12% cost reduction. Matthew LeMay, Shijian Li, Tian Guo 0001 |
IC2E | 3 |
| 2020 | MDINFERENCE: Balancing Inference Accuracy and Latency for Mobile ApplicationsabstractDeep Neural Networks are allowing mobile devices to incorporate a wide range of features into user applications. However, the computational complexity of these models makes it difficult to run them effectively on resource-constrained mobile devices. Prior work approached the problem of supporting deep learning in mobile applications by either decreasing model complexity or utilizing powerful cloud servers. These approaches each only focus on a single aspect of mobile inference and thus they often sacrifice overall performance.In this work we introduce a holistic approach to designing mobile deep inference frameworks. We first identify the key goals of accuracy and latency for mobile deep inference and the conditions that must be met to achieve them. We demonstrate our holistic approach through the design of a hypothetical framework called MDINFERENCE. This framework leverages two complementary techniques; a model selection algorithm that chooses from a set of cloud-based deep learning models to improve inference accuracy and an on-device request duplication mechanism to bound latency. Through empirically-driven simulations we show that MDINFERENCE improves aggregate accuracy over static approaches by over 40% without incurring SLA violations. Additionally, we show that with a target latency of 250ms, MDINFERENCE increased the aggregate accuracy in 99.74% cases on faster university networks and 96.84% cases on residential networks. Samuel S. Ogden, Tian Guo 0001 |
IC2E | 2 |
| 2020 | Characterizing and Modeling Distributed Training with Transient Cloud GPU ServersabstractCloud GPU servers have become the de facto way for deep learning practitioners to train complex models on large-scale datasets. However, it is challenging to determine the appropriate cluster configuration-e.g., server type and number-for different training workloads while balancing the trade-offs in training time, cost, and model accuracy. Adding to the complexity is the potential to reduce the monetary cost by using cheaper, but revocable, transient GPU servers.In this work, we analyze distributed training performance under diverse cluster configurations using CM-DARE, a cloud-based measurement and training framework. Our empirical datasets include measurements from three GPU types, six geographic regions, twenty convolutional neural networks, and thousands of Google Cloud servers. We also demonstrate the feasibility of predicting training speed and overhead using regression-based models. Finally, we discuss potential use cases of our performance modeling such as detecting and mitigating performance bottlenecks. Shijian Li, Robert J. Walls, Tian Guo 0001 |
ICDCS | 3 |
| 2020 | DistStream: An Order-Aware Distributed Framework for Online-Offline Stream Clustering AlgorithmsabstractStream clustering is an important data mining technique to capture the evolving patterns in real-time data streams. Today’s data streams, e.g., IoT events and Web clicks, are usually high-speed and contain dynamically-changing patterns. Existing stream clustering algorithms usually follow an online-offline paradigm with a one-record-at-a-time update model, which was designed for running in a single machine. These stream clustering algorithms, with this sequential update model, cannot be efficiently parallelized and fail to deliver the required high throughput for stream clustering.In this paper, we present DistStream, a distributed framework that can effectively scale out online-offline stream clustering algorithms. To parallelize these algorithms for high throughput, we develop a mini-batch update model with efficient parallelization approaches. To maintain high clustering quality, DistStream’s mini-batch update model preserves the update order in all the computation steps during parallel execution, which can reflect the recent changes for dynamically-changing streaming data. We implement DistStream atop Spark Streaming, as well as four representative stream clustering algorithms based on DistStream. Our evaluation on three real-world datasets shows that DistStream-based stream clustering algorithms can achieve sublinear throughput gain and comparable (99%) clustering quality with their single-machine counterparts. Lijie Xu, Xingtong Ye, Tian Guo 0001, Wensheng Dou, Wei Wang 0049, Jun Wei 0001 |
ICDCS | 4 |
| 2020 | Recurrent Networks for Guided Multi-Attention ClassificationabstractAttention-based image classification has gained increasing popularity in recent years. State-of-the-art methods for attention-based classification typically require a large training set and operate under the assumption that the label of an image depends solely on a single object (i.e. region of interest) in the image. However, in many real-world applications (e.g. medical imaging), it is very expensive to collect a large training set. Moreover, the label of each image is usually determined jointly by multiple regions of interest (ROIs). Fortunately, for such applications, it is often possible to collect the locations of the ROIs in each training image. In this paper, we study the problem of guided multi-attention classification, the goal of which is to achieve high accuracy under the dual constraints of (1) small sample size, and (2) multiple ROIs for each image. We propose a model, called Guided Attention Recurrent Network (GARN), for multi-attention classification. Different from existing attention-based methods, GARN utilizes guidance information regarding multiple ROIs thus allowing it to work well even when sample size is small. Empirical studies on three different visual tasks show that our guided attention approach can effectively boost model performance for multi-attention image classification. Xiangnan Kong, Tian Guo 0001, John Boaz Lee, Xinyue Liu 0003, Constance M. Moore |
KDD | 3 |
| 2020 | Grad: Learning for Overhead-aware Adaptive Video Streaming with Scalable Video CodingabstractVideo streaming commonly uses Dynamic Adaptive Streaming over HTTP (DASH) to deliver good Quality of Experience (QoE) to users. Videos used in DASH are predominantly encoded by single-layered video coding such as H.264/AVC. In comparison, multi-layered video coding such as H.264/SVC provides more flexibility for upgrading the quality of buffered video segments and has the potential to further improve QoE. However, there are two challenges for using SVC in DASH: (i) the complexity in designing ABR algorithms; and (ii) the negative impact of SVC's coding overhead. In this work, we propose a deep reinforcement learning method called Grad for designing ABR algorithms that take advantage of the quality upgrade mechanism of SVC. Additionally, we quantify the impact of coding overhead on the achievable QoE of SVC in DASH, and propose jump-enabled hybrid coding (HYBJ) to mitigate the impact. Through emulation, we demonstrate that Grad-HYBJ, an ABR algorithm for HYBJ learned by Grad, outperforms the best performing state-of-the-art ABR algorithm by 17% in QoE. Yunzhuo Liu, Bo Jiang 0003, Tian Guo 0001, Ramesh K. Sitaraman, Don Towsley, Xinbing Wang |
ACM Multimedia | 3 |
| 2020 | VVSec: Securing Volumetric Video Streaming via Benign Use of Adversarial PerturbationabstractVolumetric video (VV) streaming has drawn an increasing amount of interests recently with the rapid advancements in consumer VR/AR devices and the relevant multimedia and graphics research. While the resource and performance challenges in volumetric video streaming have been actively investigated by the multimedia community, the potential security and privacy concerns with this new type of multimedia have not been studied. We for the first time identify an effective threat model that extracts 3D face models from volumetric videos and compromises face ID-based authentications To defend against such attack, we develop a novel volumetric video security mechanism, namely VVSec, which makes benign use of adversarial perturbations to obfuscate the security and privacy-sensitive 3D face models. Such obfuscation ensures that the 3D models cannot be exploited to bypass deep learning-based face authentications. Meanwhile, the injected perturbations are not perceivable by the end-users, maintaining the original quality of experience in volumetric video streaming. We evaluate VVSec using two datasets, including a set of frames extracted from an empirical volumetric video and a public RGB-D face image dataset. Our evaluation results demonstrate the effectiveness of both the proposed attack and defense mechanisms in volumetric video streaming. Zhongze Tang, Xianglong Feng, Yi Xie 0001, Huy Phan, Tian Guo 0001, Bo Yuan 0001, Sheng Wei 0001 |
ACM Multimedia | 5 |
| 2020 | QuRate: power-efficient mobile immersive video streamingabstractSmartphones have recently become a popular platform for deploying the computation-intensive virtual reality (VR) applications, such as immersive video streaming (a.k.a., 360-degree video streaming). One specific challenge involving the smartphone-based head mounted display (HMD) is to reduce the potentially huge power consumption caused by the immersive video. To address this challenge, we first conduct an empirical power measurement study on a typical smartphone immersive streaming system, which identifies the major power consumption sources. Then, we develop QuRate, a quality-aware and user-centric frame rate adaptation mechanism to tackle the power consumption issue in immersive video streaming. QuRate optimizes the immersive video power consumption by modeling the correlation between the perceivable video quality and the user behavior. Specifically, QuRate builds on top of the user's reduced level of concentration on the video frames during view switching and dynamically adjusts the frame rate without impacting the perceivable video quality. We evaluate QuRate with a comprehensive set of experiments involving 5 smartphones, 21 users, and 6 immersive videos using empirical user head movement traces. Our experimental results demonstrate that QuRate is capable of extending the smartphone battery life by up to 1.24X while maintaining the perceivable video quality during immersive video streaming. Also, we conduct an Institutional Review Board (IRB)-approved subjective user study to further validate the minimum video quality impact caused by QuRate. Nan Jiang 0020, Yao Liu 0001, Tian Guo 0001, Wenyao Xu, Viswanathan (Vishy) Swaminathan, Lisong Xu, Sheng Wei 0001 |
MMSys | 3 |
| 2019 | An Experimental Evaluation of Garbage Collectors on Big Data ApplicationsabstractPopular big data frameworks, ranging from Hadoop MapReduce to Spark, rely on garbage-collected languages, such as Java and Scala. Big data applications are especially sensitive to the effectiveness of garbage collection (i.e., GC), because they usually process a large volume of data objects that lead to heavy GC overhead. Lacking in-depth understanding of GC performance has impeded performance improvement in big data applications. In this paper, we conduct the first comprehensive evaluation on three popular garbage collectors, i.e., Parallel, CMS, and G1, using four representative Spark applications. By thoroughly investigating the correlation between these big data applications' memory usage patterns and the collectors' GC patterns, we obtain many findings about GC inefficiencies. We further propose empirical guidelines for application developers, and insightful optimization strategies for designing big-data-friendly garbage collectors. Lijie Xu, Tian Guo 0001, Wensheng Dou, Wei Wang 0049, Jun Wei 0001 |
Proc. VLDB Endow. | 2 |
| 2018 | Cloud-Based or On-Device: An Empirical Study of Mobile Deep InferenceabstractModern mobile applications are benefiting significantly from the advancement in deep learning, e.g., implementing real-time image recognition and conversational system. Given a trained deep learning model, applications usually need to perform a series of matrix operations based on the input data, in order to infer possible output values. Because of computational complexity and size constraints, these trained models are often hosted in the cloud. To utilize these cloud-based models, mobile apps will have to send input data over the network. While cloud-based deep learning can provide reasonable response time for mobile apps, it restricts the use case scenarios, e.g. mobile apps need to have network access. With mobile specific deep learning optimizations, it is now possible to employ on-device inference. However, because mobile hardware, such as GPU and memory size, can be very limited when compared to its desktop counterpart, it is important to understand the feasibility of this new on-device deep learning inference architecture. In this paper, we empirically evaluate the inference performance of three Convolutional Neural Networks (CNNs) using a benchmark Android application we developed. Our measurement and analysis suggest that on-device inference can cost up to two orders of magnitude greater response time and energy when compared to cloud-based inference, and that loading model and computing probability are two performance bottlenecks for on-device deep inferences. Tian Guo 0001 |
IC2E | 1 |
| 2018 | Latency-aware virtual desktops optimization in distributed clouds
Tian Guo 0001, Prashant J. Shenoy, K. K. Ramakrishnan, Vijay Gopalakrishnan |
Multim. Syst. | 1 |
| 2018 | Performance and Cost Considerations for Providing Geo-Elasticity in Database CloudsabstractOnline applications that serve global workload have become a norm and those applications are experiencing not only temporal but also spatial workload variations. In addition, more applications are hosting their backend tiers separately for benefits such as ease of management. To provision for such applications, traditional elasticity approaches that only consider temporal workload dynamics and assume well-provisioned backends are insufficient. Instead, in this article, we propose a new type of provisioning mechanisms—geo-elasticity, by utilizing distributed clouds with different locations. Centered on this idea, we build a system called DBScale that tracks geographic variations in the workload to dynamically provision database replicas at different cloud locations across the globe. Our geo-elastic provisioning approach comprises a regression-based model that infers database query workload from spatially distributed front-end workload, a two-node open queueing network model that estimates the capacity of databases serving both CPU and I/O-intensive query workloads and greedy algorithms for selecting best cloud locations based on latency and cost. We implement a prototype of our DBScale system on Amazon EC2’s distributed cloud. Our experiments with our prototype show up to a 66% improvement in response time when compared to local elasticity approaches. Tian Guo 0001, Prashant J. Shenoy |
ACM Trans. Auton. Adapt. Syst. | 1 |
| 2018 | Providing Geo-Elasticity in Geographically Distributed CloudsabstractGeographically distributed cloud platforms are well suited for serving a geographically diverse user base. However, traditional cloud provisioning mechanisms that make local scaling decisions are not adequate for delivering the best possible performance for modern web applications that observe both temporal and spatial workload fluctuations. We propose GeoScale, a system that provides geo-elasticity by combining model-driven proactive and agile reactive provisioning approaches. GeoScale can dynamically provision server capacity at any location based on workload dynamics. We conduct a detailed evaluation of GeoScale on Amazon’s geo-distributed cloud and show up to 40% improvement in the 95th percentile response time when compared to traditional elasticity techniques. Tian Guo 0001, Prashant J. Shenoy |
ACM Trans. Internet Techn. | 1 |
| 2018 | Managing Risk in a Derivative IaaS CloudabstractInfrastructure-as-a-Service (IaaS) cloud platforms rent computing resources with different cost and availability tradeoffs. For example, users may acquire virtual machines (VMs) in the spot market-that are cheap, but can be unilaterally terminated by the cloud operator. Because of this revocation risk, spot servers have been conventionally used for delay and risk tolerant batch jobs. In this paper, we develop risk mitigation policies which allow even interactive applications to run on spot servers. Our System, SpotCheck is a derivative cloud platform, and provides the illusion of an IaaS platform that offers always-available VMs on demand for a cost near that of spot servers, and supports unmodified applications. SpotCheck's design combines virtualization-based mechanisms for fault-tolerance, and bidding and server selection policies for managing the risk and cost. We implement SpotCheck on EC2 and show that it i) provides nested VMs with 99.9989 percent availability, ii) achieves upto 2-5x cost savings compared to using on-demand VMs, and iii) eliminates any risk of losing VM state. Prateek Sharma 0001, Stephen Lee, Tian Guo 0001, David Irwin 0001, Prashant J. Shenoy |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2016 | Flint: batch-interactive data-intensive processing on transient serversabstractCloud providers now offer transient servers, which they may revoke at anytime, for significantly lower prices than on-demand servers, which they cannot revoke. The low price of transient servers is particularly attractive for executing an emerging class of workload, which we call Batch-Interactive Data-Intensive (BIDI), that is becoming increasingly important for data analytics. BIDI workloads require large sets of servers to cache massive datasets in memory to enable low latency operation. In this paper, we illustrate the challenges of executing BIDI workloads on transient servers, where revocations (akin to failures) are the common case. To address these challenges, we design Flint, which is based on Spark and includes automated checkpointing and server selection policies that i) support batch and interactive applications and ii) dynamically adapt to application characteristics. We evaluate a prototype of Flint using EC2 spot instances, and show that it yields cost savings of up to 90% compared to using on-demand servers, while increasing running time by < 2%. Prateek Sharma 0001, Tian Guo 0001, David Irwin 0001, Prashant J. Shenoy |
EuroSys | 2 |
| 2016 | GeoScale: Providing Geo-Elasticity in Distributed CloudsabstractDistributed cloud platforms are well suited for serving a geographically diverse user base. However traditional cloud provisioning mechanisms that make local scaling decisions are not well suited for temporal and spatial workload fluctuations seen by modern web applications. In this paper, we argue the need of geo-elasticity and present GeoScale, a system to provide geo-elasticity in distributed clouds. We describe GeoScale's model-driven proactive provisioning approach and conduct an initial evaluation of GeoScale on Amazon's distributed EC2 cloud. Our results show up to 31% improvement in the 95th percentile response time when compared to traditional elasticity techniques. Tian Guo 0001, Prashant J. Shenoy, Hakan Hacigümüs |
IC2E | 1 |
| 2016 | Analyzing the Efficiency of a Green University Data CenterabstractData centers are an indispensable part of today's IT infrastructure. To keep pace with modern computing needs, data centers continue to grow in scale and consume increasing amounts of power. While prior work on data centers has led to significant improvements in their energy-efficiency, detailed measurements from these facilities' operations are not widely available, as data center design is often considered part of a company's competitive advantage. However, such detailed measurements are critical to the research community in motivating and evaluating new energy-efficiency optimizations. In this paper, we present a detailed analysis of a state-of-the-art 15MW green multi-tenant data center that incorporates many of the technological advances used in commercial data centers. We analyze the data center's computing load and its impact on power, water, and carbon usage using standard effectiveness metrics, including PUE, WUE, and CUE. Our results reveal the benefits of optimizations, such as free cooling, and provide insights into how the various effectiveness metrics change with the seasons and increasing capacity usage. More broadly, our PUE, WUE, and CUE analysis validate the green design of this LEED Platinum data center. Patrick Pegus II, Benoy Varghese, Tian Guo 0001, David Irwin 0001, Prashant J. Shenoy, Anirban Mahanti, James Culbert, John Goodhue, Chris Hill |
ICPE | 3 |
| 2015 | SpotOn: a batch computing service for the spot marketabstractCloud spot markets enable users to bid for compute resources, such that the cloud platform may revoke them if the market price rises too high. Due to their increased risk, revocable resources in the spot market are often significantly cheaper (by as much as 10×) than the equivalent non-revocable on-demand resources. One way to mitigate spot market risk is to use various fault-tolerance mechanisms, such as checkpointing or replication, to limit the work lost on revocation. However, the additional performance overhead and cost for a particular fault-tolerance mechanism is a complex function of both an application's resource usage and the magnitude and volatility of spot market prices. Supreeth Subramanya, Tian Guo 0001, Prateek Sharma 0001, David Irwin 0001, Prashant J. Shenoy |
SoCC | 2 |
| 2015 | SpotCheck: designing a derivative IaaS cloud on the spot marketabstractInfrastructure-as-a-Service (IaaS) cloud platforms rent resources, in the form of virtual machines (VMs), under a variety of contract terms that offer different levels of risk and cost. For example, users may acquire VMs in the spot market that are often cheap but entail significant risk, since their price varies over time based on market supply and demand and they may terminate at any time if the price rises too high. Currently, users must manage all the risks associated with using spot servers. As a result, conventional wisdom holds that spot servers are only appropriate for delay-tolerant batch applications. In this paper, we propose a derivative cloud platform, called SpotCheck, that transparently manages the risks associated with using spot servers for users. Prateek Sharma 0001, Stephen Lee, Tian Guo 0001, David Irwin 0001, Prashant J. Shenoy |
EuroSys | 3 |
| 2014 | VMShadow: optimizing the performance of latency-sensitive virtual desktops in distributed cloudsabstractDistributed clouds offer a choice of data center locations to application providers to host their applications. In this paper we consider distributed clouds that host virtual desktops(VDs) which are then accessed by their users through remote desktop protocols. VDs have different sensitivities to latency, primarily determined by the types of applications running (games or video players are more sensitive to latency) and the end users' locations. We design VMShadow, a system to automatically optimize the location and performance of latency-sensitive VDs in the cloud. VMShadow performs black-box fingerprinting of a VM's network traffic to infer its latency-sensitivity and employs a greedy heuristic based algorithm to move highly latency-sensitive VMs to cloud sites that are closer to their end users. VMShadow employs WAN-based live migration and a new network connection migration protocol to ensure that the VM migration and subsequent changes to the VM's network address are transparent to end-users. We implement a prototype of VMShadow in a nested hypervisor and demonstrate its effectiveness for optimizing the performance of VM-based desktops in the cloud. Our experiments on a private and the public EC2 cloud show that VMShadow is able to discriminate between latency-sensitive and insensitive desktop applications and judiciously move only those VMs that will benefit the most. For desktop VMs with video activity, VMShadow improves VNC's refresh rate by 90%. Further our connection migration proxy, which utilizes dynamic rewriting of packet headers, imposes a rewriting overhead of only 13μs per packet. Trans-continental VM migrations take about 4 minutes. Tian Guo 0001, Vijay Gopalakrishnan, K. K. Ramakrishnan, Prashant J. Shenoy, Arun Venkataramani, Seungjoon Lee |
MMSys | 1 |
| 2014 | Cost-Aware Cloud Bursting for Enterprise ApplicationsabstractThe high cost of provisioning resources to meet peak application demands has led to the widespread adoption of pay-as-you-go cloud computing services to handle workload fluctuations. Some enterprises with existing IT infrastructure employ a hybrid cloud model where the enterprise uses its own private resources for the majority of its computing, but then “bursts” into the cloud when local resources are insufficient. However, current commercial tools rely heavily on the system administrator’s knowledge to answer key questions such as when a cloud burst is needed and which applications must be moved to the cloud. In this article, we describe Seagull, a system designed to facilitate cloud bursting by determining which applications should be transitioned into the cloud and automating the movement process at the proper time. Seagull optimizes the bursting of applications using an optimization algorithm as well as a more efficient but approximate greedy heuristic. Seagull also optimizes the overhead of deploying applications into the cloud using an intelligent precopying mechanism that proactively replicates virtualized applications, lowering the bursting time from hours to minutes. Our evaluation shows over 100% improvement compared to naïve solutions but produces more expensive solutions compared to ILP. However, the scalability of our greedy algorithm is dramatically better as the number of VMs increase. Our evaluation illustrates scenarios where our prototype can reduce cloud costs by more than 45% when bursting to the cloud, and that the incremental cost added by precopying applications is offset by a burst time reduction of nearly 95%. Tian Guo 0001, Upendra Sharma, Prashant J. Shenoy, Timothy Wood 0001, Sambit Sahu |
ACM Trans. Internet Techn. | 1 |
| 2013 | VMShadow: optimizing the performance of virtual desktops in distributed cloudsabstractWe present VMShadow, a system that automatically optimizes the location and performance of applications based on their dynamic workloads. We prototype VMShadow and demonstrate its efficacy using VM-based desktops in the cloud as an example application. Our experiments on a private cloud as well as the EC2 cloud, using a nested hypervisor, show that VMShadow is able to discriminate between location-sensitive and location-insensitive desktop VMs and judiciously moves only those that will benefit the most from the migration. For example, VMShadow performs transcontinental VM migrations in ~ 4 mins and can improve VNC's video refresh rate by up to 90%. Tian Guo 0001, Vijay Gopalakrishnan, K. K. Ramakrishnan, Prashant J. Shenoy, Arun Venkataramani, Seungjoon Lee |
SoCC | 1 |
| 2012 | Seagull: Intelligent Cloud Bursting for Enterprise Applications
Tian Guo 0001, Upendra Sharma, Timothy Wood 0001, Sambit Sahu, Prashant J. Shenoy |
USENIX ATC | 1 |