Timothy Wood 0001

dblp:w/TimothyWood · also Tim Wood 0006 · DBLP profile ↗
← Back
58ranked-venue papers
9as first author
9since 2021 · last 2025
0000-0002-6728-4197ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 4 first-author · 6 since 2021Computer networks · 18 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 first-authorSecurity and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Poster: "Semantic-Guided Skip Sampling and Soft Edge Compression for Real-Time Edge-Based Traffic Video Analytics"
abstract
Edge-based traffic monitoring is an important component in building smart transportation infrastructure that understands the flow of vehicles and pedestrians. Yet the necessary multi-modal data including video, in-vehicle sensors, and embedded traffic monitors result in large-scale continuous data streams that are difficult to process in real-time on edge devices and would be wasteful to send to a distant cloud. In this paper, we propose an integrated framework combining semantically-guided skip sampling with soft edge-based neural compression to enable efficient traffic video analytics on resource-constrained edge devices.
Deng Pan 0010, Timothy Wood 0001, Samer Hamdar
SEC2
2025 SledgeScale: Load-Aware Dispatch and Deadline-Driven Scheduling for Scalable, Dense Serverless Computing in Edge Data Centers
abstract
Serverless and Edge Computing ought to be a perfect match to create flexible, efficient, and responsive applications for emerging areas like augmented reality and autonomous vehicles. Unfortunately, current serverless designs incur high overheads—preventing submillisecond execution—and consume large amounts of resources—preventing dense deployment within constrained edge environments. We present a platform to overcome these challenges, while providing strong isolation and performance in multi-tenant edge data centers.
Xiaosu Lyu, Emil Abbasov, Sean McBride, Gabriel Parmer, Timothy Wood 0001
SEC5
2024 Byways: High-Performance, Isolated Network Functions for Multi-Tenant Cloud Servers
abstract
Network functions (NFs) have become pervasive in data centers as a means to monitor and transform traffic as it flows between services. Softwarization of the network has further added to the diversity of functions that can be deployed, yet managing the performance, efficiency, tenant-customizability, and security of these functions remains a major challenge. We present Byways, an abstraction that provides facilitates to safely deploy NFs alongside end-host VMs in a multi-tenant cloud environment. Byways guarantee strict isolation between the host system, the network functions, and VM-based cloud applications, while still maintaining high performance. A Byway manages a specific set of services, and an associated NF only processes flows associated with those services, using per-byway resources (e.g., processing time). This separation of end-host traffic across Byways provides strong fault isolation - a failing NF does not impact other services. Byways augment this isolation with per-Byway access rights that restrict a NFs access (e.g., read, write, drop) to the flow, limiting the impact of a faulty NF on even its own services.
Yuan Gao 0041, Gabriel Parmer, Timothy Wood 0001
SoCC4
2024 Envisioning a Unified Programmable Dataplane to Monitor Slow Attacks
abstract
Recent work shows that programmable switches can effectively detect attack traffic, such as denial-of-service attacks in the midst of high-volume network traffic. However, these techniques primarily rely on sampling or sketch-based data structures, which can only be used to approximate the characteristics of dominant flows in the network. As a result, such techniques are unable to effectively detect low-volume attacks that stealthily add only a few packets to the network. Our work explores how the combination of programmable switches, Smart network interface cards, and hosts can enable fine-grained analysis of every flow in a network, even those with only a small number of packets. We focus on analyzing packets at the start of each flow, as those packets often can help indicate whether a flow is benign or suspicious. We propose a unified architecture that spans the full programmable dataplane to take advantage of the strengths of each type of device. We are developing new filter data structures to efficiently track flows on the switch, dataplane-based communication protocols to quickly coordinate between devices, and caching approaches on the SmartNIC that help minimize the traffic load reaching the host. Our preliminary prototype can handle the full pipe bandwidth of 1.4 Tbps of traffic entering the Tofino switch, forward only 20 Gbps to the SmartNIC, and minimize the traffic load to 5 Gbps reaching the host due to our efficient flow filter, packet batching, and SmartNIC-based cache.
Cuidi Wei, Shaoyu Tu, Toru Hasegawa, Yuki Koizumi, K. K. Ramakrishnan, Junji Takemasa, Timothy Wood 0001
ICNP7
2024 Dynamically Balancing Load with Overload Control for Microservices
abstract
The microservices architecture simplifies application development by breaking monolithic applications into manageable microservices. However, this distributed microservice “service mesh” leads to new challenges due to the more complex application topology. Particularly, each service component scales up and down independently creating load imbalance problems on shared backend services accessed by multiple components. Traditional load balancing algorithms do not port over well to a distributed microservice architecture where load balancers are deployed client-side. In this article, we propose a self-managing load balancing system, BLOC, which provides consistent response times to users without using a centralized metadata store or explicit messaging between nodes. BLOC uses overload control approaches to provide feedback to the load balancers. We show that this performs significantly better in solving the incast problem in microservice architectures. A critical component of BLOC is the dynamic capacity estimation algorithm. We show that a well-tuned capacity estimate can outperform even join-the-shortest-queue, a nearly optimal algorithm, while a reasonable dynamic estimate still outperforms Least Connection, a distributed implementation of join-the-shortest-queue. Evaluating this framework, we found that BLOC improves the response time distribution range, between the 10th and 90th percentiles, by 2 –4 times and the tail, 99th percentile, latency by 2 times.
Ratnadeep Bhattacharya, Yuan Gao 0041, Timothy Wood 0001
ACM Trans. Auton. Adapt. Syst.3
2023 Sidecar-based Path-aware Security for Microservices
abstract
Microservice architectures decompose web applications into loosely-coupled, distributed components that interact with each other to provide an overall service. While this popular software architecture paradigm has many advantages in development and deployment, it also introduces a wider attack surface that is vulnerable to both internal and external attackers. Potentially malicious third-party services or software packages, as well as increased communication endpoints, introduce a wide array of security concerns. To improve the resiliency of microservice-based applications, many of which store sensitive data, we propose a novel, path-based anomaly detection and access control infrastructure that requires no modifications to existing software. We propose leveraging trusted proxies deployed alongside each service for request inspection, anomaly detection and signed token propagation for end-user path validation. Our approach reduces the trusted computing base away from the microservices to a smaller set of components that allow for less trust and a smaller attack surface.
Catherine Meadows 0002, Sena Hounsinou, Timothy Wood 0001, Gedare Bloom
SACMAT3
2022 Poster: Toward Zero-Trust Path-Aware Access Control
abstract
In this poster, we introduce path-aware risk scores for access control (PARSAC), a novel context-sensitive technique to enrich access requests with risk scoring of the path taken by those requests between the authenticated user and the resources they access. These path-aware risk scores enable another layer of security for traditional access control systems that addresses the need for fine-grained monitoring and enforcement within a zero-trust architecture. We define rules for general functions that can be used to determine risk and instantiate a specific approach to calculate path risk scores. We evaluate our approach with realistic network graphs; PARSAC finds more paths with lower risk when compared with traditional routing algorithms that select the shortest path.
Joshua H. Seaton, Sena Hounsinou, Timothy Wood 0001, Shouhuai Xu, Philip N. Brown, Gedare Bloom
SACMAT3
2021 Mu: An Efficient, Fair and Responsive Serverless Framework for Resource-Constrained Edge Clouds
abstract
Serverless computing platforms simplify development, deployment, and automated management of modular software functions. However, existing serverless platforms typically assume an over-provisioned cloud, making them a poor fit for Edge Computing environments where resources are scarce. In this paper we propose a redesigned serverless platform that comprehensively tackles the key challenges for serverless functions in a resource constrained Edge Cloud.
Viyom Mittal, Shixiong Qi, Ratnadeep Bhattacharya, Xiaosu Lyu, Sameer G. Kulkarni, Dan Li 0001, Jinho Hwang, K. K. Ramakrishnan, Timothy Wood 0001
SoCC10
2021 OFC: an opportunistic caching system for FaaS platforms
abstract
Cloud applications based on the "Functions as a Service" (FaaS) paradigm have become very popular. Yet, due to their stateless nature, they must frequently interact with an external data store, which limits their performance. To mitigate this issue, we introduce OFC, a transparent, vertically and horizontally elastic in-memory caching system for FaaS platforms, distributed over the worker nodes. OFC provides these benefits cost-effectively by exploiting two common sources of resource waste: (i) most cloud tenants overprovision the memory resources reserved for their functions because their footprint is non-trivially input-dependent and (ii) FaaS providers keep function sandboxes alive for several minutes to avoid cold starts. Using machine learning models adjusted for typical function input data categories (e.g., multimedia formats), OFC estimates the actual memory resources required by each function invocation and hoards the remaining capacity to feed the cache. We build our OFC prototype based on enhancements to the OpenWhisk FaaS platform, the Swift persistent object store, and the RAM-Cloud in-memory store. Using a diverse set of workloads, we show that OFC improves by up to 82 % and 60 % respectively the execution time of single-stage and pipelined functions.
Djob Mvondo, Mathieu Bacou, Kevin Nguetchouang, Lucien Ngale, Stéphane Pouget, Josiane Kouam, Renaud Lachaize, Jinho Hwang, Timothy Wood 0001, Daniel Hagimont, Noel De Palma, Bernabe Batchakui, Alain Tchana
EuroSys9
2020 Managing State for Failure Resiliency in Network Function Virtualization
abstract
Ensuring high scalability (elastic scale-out and consolidation), as well as high availability (failure resiliency) are critical in encouraging adoption of software-based network functions (NFs). In recent years, two paradigms have evolved in terms of the way the NFs manage their state - namely the Stateful (state is coupled with the NF instance) and a Stateless (state is externalized to a datastore) manner. These two paradigms present unique challenges and opportunities for ensuring high scalability and high availability of NFs and NF chains. In this work, we assess the impact on ensuring the correctness of NF state including the implications of non-determinism in packet processing, and carefully analyze and present the benefits and disadvantages of the two state management paradigms. We leverage OpenNetVM and Redis in-memory datastore to implement both state management paradigms and empirically compare the two. Although the stateless paradigm is desirable for elastic scaling, our experimental results show that, even at line-rate packet processing (10 Gbps), stateful NFs can achieve chain-level failover across servers in a LAN incurring less than 10% performance. The state-of-the-art stateless counterparts incur severe throughput penalties. We observe 30-85% overhead on normal processing, depending on the mode of state updated to the externalized datastore.
Sameer G. Kulkarni, K. K. Ramakrishnan, Timothy Wood 0001
LANMAN3
2020 Fine-Grained Isolation for Scalable, Dynamic, Multi-tenant Edge Clouds
Yuxin Ren 0001, Guyue Liu, Vlad Nitu, Wenyuan Shao, Riley Kennedy, Gabriel Parmer, Timothy Wood 0001, Alain Tchana
USENIX ATC7
2020 Introduction to the Special Issue with Selected Papers of The International Conference on Autonomic Computing and Self-Organizing Systems (ACSOS) 2020
abstract
No abstract available.
Sven Tomforde, Timothy Wood 0001, Jan-Philipp Steghöfer
ACM Trans. Auton. Adapt. Syst.2
2020 REINFORCE: Achieving Efficient Failure Resiliency for Network Function Virtualization-Based Services
abstract
Ensuring high availability (HA) for software-based networks is a critical design feature that will help the adoption of software-based network functions (NFs) in production networks. It is important for NFs to avoid outages and maintain mission-critical operations. However, HA support for NFs on the critical data path can result in unacceptable performance degradation. We present REINFORCE, an integrated framework to support efficient resiliency for NF service chains. REINFORCE includes timely failure detection and consistent failover mechanisms. REINFORCE replicates state to standby NFs (local and remote) while enforcing correctness. It minimizes the number of state transfers by exploiting the concept of external synchrony, and leverages opportunistic batching and multi-buffering to optimize performance. Experimental results show that, even at line-rate packet processing (10 Gbps), REINFORCE achieves chain-level failover across servers in a LAN within 10ms, incurring less than 10% performance overhead, and adds average latency only ~400 μs, with a worst-case latency of less than 1ms. REINFORCE also recovers from software failures within the same node in less than 100 μs, incurring less than 1% performance overhead and adds less than 5 μs latency during normal operation.
Sameer G. Kulkarni, Guyue Liu, K. K. Ramakrishnan, Mayutan Arumaithurai, Timothy Wood 0001, Xiaoming Fu 0001
IEEE/ACM Trans. Netw.5
2020 NFVnice: Dynamic Backpressure and Scheduling for NFV Service Chains
abstract
Managing Network Function (NF) service chains requires careful system resource management. We propose NFVnice, a user space NF scheduling and service chain management framework to provide fair, efficient and dynamic resource scheduling capabilities on Network Function Virtualization (NFV) platforms. The NFVnice framework monitors load on a service chain at high frequency (1000Hz) and employs backpressure to shed load early in the service chain, thereby preventing wasted work. Borrowing concepts such as rate proportional scheduling from hardware packet schedulers, CPU shares are computed by accounting for heterogeneous packet processing costs of NFs, I/O, and traffic arrival characteristics. By leveraging cgroups, a user space process scheduling abstraction exposed by the operating system, NFVnice is capable of controlling when network functions should be scheduled. NFVnice improves NF performance by complementing the capabilities of the OS scheduler but without requiring changes to the OS's scheduling mechanisms. Our controlled experiments show that NFVnice provides the appropriate rate-cost proportional fair share of CPU to NFs and significantly improves NF performance (throughput and latency) by reducing wasted work across an NF chain, compared to using the default OS scheduler. NFVnice achieves this even for heterogeneous NFs with vastly different computational costs and for heterogeneous workloads.
Sameer G. Kulkarni, Wei Zhang 0052, Jinho Hwang, Shriram Rajagopalan, K. K. Ramakrishnan, Timothy Wood 0001, Mayutan Arumaithurai, Xiaoming Fu 0001
IEEE/ACM Trans. Netw.6
2019 Living on the Edge: Serverless Computing and the Cost of Failure Resiliency
abstract
Serverless computing platforms have gained popularity because they allow easy deployment of services in a highly scalable and cost-effective manner. By enabling just-in-time startup of container-based services, these platforms can achieve good multiplexing and automatically respond to traffic growth, making them particularly desirable for edge cloud data centers where resources are scarce. Edge cloud data centers are also gaining attention because of their promise to provide responsive, low-latency shared computing and storage resources. Bringing serverless capabilities to edge cloud data centers must continue to achieve the goals of low latency and reliability. The reliability guarantees provided by serverless computing however are weak, with node failures causing requests to be dropped or executed multiple times. Thus serverless computing only provides a best effort infrastructure, leaving application developers responsible for implementing stronger reliability guarantees at a higher level. Current approaches for providing stronger semantics such as “exactly once” guarantees could be integrated into serverless platforms, but they come at high cost in terms of both latency and resource consumption. As edge cloud services move towards applications such as autonomous vehicle control that require strong guarantees for both reliability and performance, these approaches may no longer be sufficient. In this paper we evaluate the latency, throughput, and resource costs of providing different reliability guarantees, with a focus on these emerging edge cloud platforms and applications.
Sameer G. Kulkarni, Guyue Liu, K. K. Ramakrishnan, Timothy Wood 0001
LANMAN4
2019 Advancing Network Function Virtualization Platforms with Programmable NICs
abstract
Network Function Virtualization seeks to run high performance middleboxes in a flexible, more configurable software environment. Even with advances such as kernel bypass and zero-copy IO, middlebox platforms still struggle to meet stringent throughput and latency requirements. To achieve line rates as network bandwidths rise, these platforms often must make tradeoffs such as inefficiently dedicating more CPU cores or weakening security and isolation properties. In this paper we explore how advances in programmable “smart NICs” can be leveraged by software middlebox platforms to improve performance, resource efficiency, and security. Our evaluation shows several use cases for smart NICs, which improve performance significantly while reducing resource consumption and providing strong isolation.
Zhen Ni, Guyue Liu, Dennis Afanasev, Timothy Wood 0001, Jinho Hwang
LANMAN4
2018 REINFORCE: achieving efficient failure resiliency for network function virtualization based services
abstract
Ensuring high availability (HA) for software-based networks is a critical design feature that will help the adoption of software-based network functions (NFs) in production networks. It is important for NFs to avoid outages and maintain mission-critical operations. However, HA support for NFs on the critical data path can result in unacceptable performance degradation. We present REINFORCE, an integrated framework to support efficient resiliency for NFs and NF service chains. REINFORCE includes timely failure detection and consistent failover mechanisms. REINFORCE replicates state to standby NFs (local and remote) while enforcing correctness. It minimizes the number of state transfers by exploiting the concept of external synchrony, and leverages opportunistic batching and multi-buffering to optimize performance. Experimental results show that, even at line-rate packet processing (10 Gbps), REINFORCE achieves chain-level failover across servers in a LAN (or within the same node) within 10ms (100/μs), incurring less than 10% (1%) performance overhead, and adds average latency of only ~400/μs (5/μs), with a worst-case latency of less than 1ms (10/μs).
Sameer G. Kulkarni, Guyue Liu, K. K. Ramakrishnan, Mayutan Arumaithurai, Timothy Wood 0001, Xiaoming Fu 0001
CoNEXT5
2018 Measuring Performance and Isolation Tradeoffs for NFV
abstract
Network function virtualization (NFV) allows network services, such as firewalls and routing, to be deployed into a virtual environment and run on commodity hardware. Recently service providers and developers can deploy their network functions (NF) prototypes on a shared infrastructure, and all the NFs are being controlled by the manager of the platform. NFV platforms run these NFs together, and share the system resources to optimize the utilization. This means that limited resources such as CPU cores or memory have to be shared. Several recent NFV systems run network services with one shared memory region, so that they can achieve high performance with zero-copy I/O. This resource sharing brings security problems since it allows malicious NFs to easily modify data from other NFs. To enhance the security of NFV, we are designing a platform to provide stronger memory isolation between different NFs. Our approach is based on the architecture developed for our OpenNetVM platform, which supports lightweight NFs, flexible management, but assumes a single shared memory pool for all NFs.
Zhen Ni, Timothy Wood 0001
LANMAN2
2018 CRIMES: Using Evidence to Secure the Cloud
abstract
Cloud applications are appealing targets to attackers, yet current cloud infrastructures have few ways of helping defend their customers from attacks. However, the use of virtual machines, and the economy of scale found in cloud platforms, provides an opportunity to offer strong security guarantees to tenants at low cost to the cloud provider. We present CRIMES, an evidence based, modular security framework for cloud platforms that uses speculative execution coupled with memory introspection tools to detect malicious behavior in real time. By buffering VM outputs (i.e., outgoing network packets and disk writes) until a scan has been completed, CRIMES gives strong guarantees about the amount of damage an attack can do, while minimizing overheads. When an attack is detected, CRIMES rolls back to a recent checkpoint and performs automated forensic analysis to help pinpoint the source of an attack. Our evaluation demonstrates that CRIMES incurs less overhead compared to memory protection tools such as AddressSanitizer, while offering valuable forensic analysis for buffer overflow attacks and malware detection across multiple applications and the OS.
Sundaresan Rajasekaran, Harpreet Singh Chawla, Zhen Ni, Emery D. Berger, Timothy Wood 0001
Middleware6
2018 Microboxes: high performance NFV with customizable, asynchronous TCP stacks and dynamic subscriptions
abstract
Existing network service chaining frameworks are based on a "packet-centric" model where each NF in a chain is given every packet for processing. This approach becomes both inefficient and inconvenient for more complex network functions that operate at higher levels of the protocol stack. We propose Microboxes, a novel service chaining abstraction designed to support transport- and application-layer middle-boxes, or even end-system like services. Simply including a TCP stack in an NFV platform is insufficient because there is a wide spectrum of middlebox types-from NFs requiring only simple TCP bytestream reconstruction to full endpoint termination. By exposing a publish/subscribe-based API for NFs to access packets or protocol events as needed, Microboxes eliminates redundant processing across a chain and enables a modular design. Our implementation on a DPDK-based NFV framework can double throughput by consolidating stack operations and provide a 51% throughput gain by customizing TCP processing to the appropriate level.
Guyue Liu, Yuxin Ren 0001, Mykola Yurchenko, K. K. Ramakrishnan, Timothy Wood 0001
SIGCOMM5
2017 Message from TPC chairs
abstract
We are pleased to welcome you to Osaka, Japan to attend IEEE LANMAN 2017, the 23rd IEEE International Symposium on Local and Metropolitan Area Networks. LANMAN began as a workshop focused on local networking technologies, but over more than two decades it has grown into a full fledged symposium covering a broad spectrum of networking issues. This year's technical program covers both wired and wireless networks, with application areas ranging from the Internet of Things, to software defined infrastructures, to cellular networks.
Tommaso Melodia, Timothy Wood 0001
LANMAN2
2017 NFVnice: Dynamic Backpressure and Scheduling for NFV Service Chains
abstract
Managing Network Function (NF) service chains requires careful system resource management. We propose NFVnice, a user space NF scheduling and service chain management framework to provide fair, efficient and dynamic resource scheduling capabilities on Network Function Virtualization (NFV) platforms. The NFVnice framework monitors load on a service chain at high frequency (1000Hz) and employs backpressure to shed load early in the service chain, thereby preventing wasted work. Borrowing concepts such as rate proportional scheduling from hardware packet schedulers, CPU shares are computed by accounting for heterogeneous packet processing costs of NFs, I/O, and traffic arrival characteristics. By leveraging cgroups, a user space process scheduling abstraction exposed by the operating system, NFVnice is capable of controlling when network functions should be scheduled. NFVnice improves NF performance by complementing the capabilities of the OS scheduler but without requiring changes to the OS's scheduling mechanisms. Our controlled experiments show that NFVnice provides the appropriate rate-cost proportional fair share of CPU to NFs and significantly improves NF performance (throughput and loss) by reducing wasted work across an NF chain, compared to using the default OS scheduler. NFVnice achieves this even for heterogeneous NFs with vastly different computational costs and for heterogeneous workloads.
Sameer G. Kulkarni, Wei Zhang 0052, Jinho Hwang, Shriram Rajagopalan, K. K. Ramakrishnan, Timothy Wood 0001, Mayutan Arumaithurai, Xiaoming Fu 0001
SIGCOMM6
2016 Flurries: Countless Fine-Grained NFs for Flexible Per-Flow Customization
abstract
The combination of Network Function Virtualization (NFV) and Software Defined Networking (SDN) allows flows to be flexibly steered through efficient processing pipelines. As deployment of NFV becomes more prevalent, the need to provide fine-grained customization of service chains and flow-level performance guarantees will increase, even as the diversity of Network Functions (NFs) rises. Existing NFV approaches typically route wide classes of traffic through pre-configured service chains. While this aggregation improves efficiency, it prevents flexibly steering and managing performance of flows at a fine granularity.
Wei Zhang 0052, Jinho Hwang, Shriram Rajagopalan, K. K. Ramakrishnan, Timothy Wood 0001
CoNEXT5
2016 Multi-cache: Dynamic, Efficient Partitioning for Multi-tier Caches in Consolidated VM Environments
abstract
Every physical machine in today's typical datacenter is backed by storage devices with hundreds of Gigabytes to Terabytes in size. Data center vendors usually use hard disk drives for their back-end storage as it is cheap and reliable. However, the increase in the I/O accesses to the back-end storage from one or many of the VMs hosted on a physical machine can reduce its overall accesses time significantly due to contention. This may not be suitable for interactive applications requiring low latency that might be co-located with other I/O intensive applications. In this paper we present Multi-Cache, a multi-layer cache management system that uses a combination of cache devices of varied speed and cost such as solid state drives, non-volatile memories, etc to mitigate this problem. Multi-Cache partitions each device dynamically at runtime according to the workload of each VM and its priority. We use a heuristic optimization technique that ensures maximum utilization of the caches resulting in a high hit rate. We use a weighted partitioning policy that improves latency by up to 72% for individual workloads, and a overall hit rate increase of up to 31% for host running several workloads together in comparison to standard LRU caching algorithms.
Sundaresan Rajasekaran, Shaohua Duan, Wei Zhang 0052, Timothy Wood 0001
IC2E4
2016 Toward online virtual network function placement in Software Defined Networks
abstract
Network function virtualization (NFV) and Software Defined Networks (SDN) separate and abstract network functions from underlying hardware, creating a flexible virtual networking environment that reduces cost and allows policy-based decisions. One of the biggest challenges in NFV-SDN is to map the required virtual network functions (VNFs) to the underlying hardware in substrate networks in a timely manner. In this paper, we formulate the VNF placement problem via Graph Pattern Matching, with an objective function that can be easily adapted to fit various applications. Previous work only considers off-line VNF placement as it is time consuming to find an appropriate mapping path while considering all software and hardware constraints. To reduce this time, we investigate the feasibility and effectiveness of path-precomputing, where paths are calculated prior to placement. Our approach enables online VNF placement in SDNs, allowing VNF requests to be processed as they arrive. An online placement approach (OPA) is proposed to place VNF requests on substrate networks. To the best of our knowledge, this is the first work in the literature that considers the online chaining VNF placement in SDNs. In addition, we present an application of OPA over cost minimization. Simulation results demonstrate that our online approach provides competitive performance compared with off-line algorithms.
Bowu Zhang, Jinho Hwang, Timothy Wood 0001
IWQoS3
2016 OpenNetVM: Flexible, high performance NFV (Demo)
abstract
Network Function Virtualization promises to enable dynamic management of software-based network functions. We envision a dynamic and flexible network that can support a smarter data plane than just simple switches that forward packets. This network architecture supports complex stateful routing of flows where processing by network functions (NFs) can transform packet data, customized on a per-flow basis, as it moves between end points. This demo will present OpenNetVM, a highly efficient packet processing framework that greatly simplifies the development of network functions, as well as their management and optimization. OpenNetVM runs network functions in lightweight Docker containers that start in less than a second. The OpenNetVM platform manager provides load balancing, flexible flow management, and service name abstractions. OpenNetVM uses DPDK for high performance I/O, and efficiently routes packets through dynamically created service chains. We will demonstrate how the research community can easily build new network functions and rapidly deploy them to see their effectiveness in high performance network environments.
Wei Zhang 0052, Guyue Liu, Phil Lopreiato, Grégoire Todeschi, K. K. Ramakrishnan, Timothy Wood 0001
LANMAN8
2016 NetAlytics: Cloud-Scale Application Performance Monitoring with SDN and NFV
Guyue Liu, Michael Trotter, Yuxin Ren 0001, Timothy Wood 0001
Middleware4
2016 SDNFV: Flexible and Dynamic Software Defined Control of an Application- and Flow-Aware Data Plane
Wei Zhang 0052, Guyue Liu, Ali Mohammadkhan, Jinho Hwang, K. K. Ramakrishnan, Timothy Wood 0001
Middleware6
2015 Cloud-Scale Application Performance Monitoring with SDN and NFV
abstract
In cloud data centers, more and more services are deployed across multiple tiers to increase flexibility and scalability. However, this makes it difficult for the cloud provider to identify which tier of the application is the bottleneck and how to resolve performance problems. Existing solutions approach this problem by constantly monitoring either in end-hosts or physical switches. Host based monitoring usually needs instrumentation of application code, making it less practical, while network hardware based monitoring is expensive and requires special features in each physical switch. Instead, we believe network wide monitoring should be flexible and easy to deploy in a non-intrusive way by exploiting recent advances in software-based network services. Towards this end we are developing a distributed software-based network monitoring framework for cloud data centers. Our system leverages knowledge of topology and routing information to build relationships between each tier of the application, and detect and locate performance bottlenecks by monitoring the network inside software switches.
Guyue Liu, Timothy Wood 0001
IC2E2
2015 Virtual function placement and traffic steering in flexible and dynamic software defined networks
abstract
The integration of network function virtualization (NFV) and software defined networks (SDN) seeks to create a more flexible and dynamic software-based network environment. The line between entities involved in forwarding and those involved in more complex middle box functionality in the network is blurred by the use of high-performance virtualized platforms capable of performing these functions. A key problem is how and where network functions should be placed in the network and how traffic is routed through them. An efficient placement and appropriate routing increases system capacity while also minimizing the delay seen by flows. In this paper, we formulate the problem of network function placement and routing as a mixed integer linear programming problem. This formulation not only determines the placement of services and routing of the flows, but also seeks to minimize the resource utilization. We develop heuristics to solve the problem incrementally, allowing us to support a large number of flows and to solve the problem for incoming flows without impacting existing flows.
Ali Mohammadkhan, Sheida Ghapani, Guyue Liu, Wei Zhang 0052, K. K. Ramakrishnan, Timothy Wood 0001
LANMAN6
2015 IOrchestra: supporting high-performance data-intensive applications in the cloud via collaborative virtualization
abstract
Multi-tier data-intensive applications are widely deployed in virtualized data centers for high scalability and reliability. As the response time is vital for user satisfaction, this requires achieving good performance at each tier of the applications in order to minimize the overall latency. However, in such virtualized environments, each tier (e.g., application, database, web) is likely to be hosted by different virtual machines (VMs) on multiple physical servers, where a guest VM is unaware of changes outside its domain, and the hypervisor also does not know the configuration and runtime status of a guest VM. As a result, isolated virtualization domains lend themselves to performance unpredictability and variance. In this paper, we propose IOrchestra, a holistic collaborative virtualization framework, which bridges the semantic gaps of I/O stacks and system information across multiple VMs, improves virtual I/O performance through collaboration from guest domains, and increases resource utilization in data centers. We present several case studies to demonstrate that IOrchestra is able to address numerous drawbacks of the current practice and improve the I/O latency of various distributed cloud applications by up to 31%.
Ron Chi-Lung Chiang, H. Howie Huang, Timothy Wood 0001, Changbin Liu, Oliver Spatscheck
SC3
2015 NetVM: High Performance and Flexible Networking Using Virtualization on Commodity Platforms
abstract
NetVM brings virtualization to the Network by enabling high bandwidth network functions to operate at near line speed, while taking advantage of the flexibility and customization of low cost commodity servers. NetVM allows customizable data plane processing capabilities such as firewalls, proxies, and routers to be embedded within virtual machines, complementing the control plane capabilities of Software Defined Networking. NetVM makes it easy to dynamically scale, deploy, and reprogram network functions. This provides far greater flexibility than existing purpose-built, sometimes proprietary hardware, while still allowing complex policies and full packet inspection to determine subsequent processing. It does so with dramatically higher throughput than existing software router platforms. NetVM is built on top of the KVM platform and Intel DPDK library. We detail many of the challenges we have solved such as adding support for high-speed inter-VM communication through shared huge pages and enhancing the CPU scheduler to prevent overheads caused by inter-core communication and context switching. NetVM allows true zero-copy delivery of data to VMs both for packet processing and messaging among VMs within a trust boundary. Our evaluation shows how NetVM can compose complex network functionality from multiple pipelined VMs and still obtain throughputs up to 10 Gbps, an improvement of more than 250% compared to existing techniques that use SR-IOV for virtualized networking.
Jinho Hwang, K. K. Ramakrishnan, Timothy Wood 0001
IEEE Trans. Netw. Serv. Manag.3
2015 CloudNet: Dynamic Pooling of Cloud Resources by Live WAN Migration of Virtual Machines
abstract
Virtualization technology and the ease with which virtual machines (VMs) can be migrated within the LAN have changed the scope of resource management from allocating resources on a single server to manipulating pools of resources within a data center. We expect WAN migration of virtual machines to likewise transform the scope of provisioning resources from a single data center to multiple data centers spread across the country or around the world. In this paper, we present the CloudNet architecture consisting of cloud computing platforms linked with a virtual private network (VPN)-based network infrastructure to provide seamless and secure connectivity between enterprise and cloud data center sites. To realize our vision of efficiently pooling geographically distributed data center resources, CloudNet provides optimized support for live WAN migration of virtual machines. Specifically, we present a set of optimizations that minimize the cost of transferring storage and virtual machine memory during migrations over low bandwidth and high-latency Internet links. We evaluate our system on an operational cloud platform distributed across the continental US. During simultaneous migrations of four VMs between data centers in Texas and Illinois, CloudNet's optimizations reduce memory migration time by 65% and lower bandwidth consumption for the storage and memory transfer by 19 GB, a 50% reduction.
Timothy Wood 0001, K. K. Ramakrishnan, Prashant J. Shenoy, Jacobus E. van der Merwe, Jinho Hwang, Guyue Liu, Lucas Chaufournier
IEEE/ACM Trans. Netw.1
2014 UniCache: Hypervisor Managed Data Storage in RAM and Flash
abstract
Application and OS-level caches are crucial for hiding I/O latency and improving application performance. However, caches are designed to greedily consume memory, which can cause memory-hogging problems in a virtualized data centers since the hypervisor cannot tell for what a virtual machine uses its memory. A group of virtual machines may contain a wide range of caches: database query pools, memcached key-value stores, disk caches, etc., each of which would like as much memory as possible. The relative importance of these caches can vary significantly, yet system administrators currently have no easy way to dynamically manage the resources assigned to a range of virtual machine data caches in a unified way. To improve this situation, we have developed UniCache, a system that provides a hypervisor managed volatile data store that can cache data either in hypervisor controlled main memory (hot data) or on Flash based storage (cold data). We propose a two-level cache management system that uses a combination of recency information, object size, and a prediction of the cost to recover an object to guide its eviction algorithm. We have built a prototype of UniCache using Xen, and have evaluated its effectiveness in a shared environment where multiple virtual machines compete for storage resources.
Jinho Hwang, Wei Zhang 0052, Ron Chi-Lung Chiang, Timothy Wood 0001, H. Howie Huang
IEEE CLOUD4
2014 MIMP: Deadline and Interference Aware Scheduling of Hadoop Virtual Machines
abstract
Virtualization promised to dramatically increase server utilization levels, yet many data centers are still only lightly loaded. In some ways, big data applications are an ideal fit for using this residual capacity to perform meaningful work, but the high level of interference between interactive and batch processing workloads currently prevents this from being a practical solution in virtualized environments. Further, the variable nature of spare capacity may make it difficult to meet big data application deadlines. In this work we propose two schedulers: one in the virtualization layer designed to minimize interference on high priority interactive services, and one in the Hadoop framework that helps batch processing jobs meet their own performance deadlines. Our approach uses performance models to match Hadoop tasks to the servers that will benefit them the most, and deadline-aware scheduling to effectively order incoming jobs. The combination of these schedulers allows data center administrators to safely mix resource intensive Hadoop jobs with latency sensitive web applications, and still achieve predictable performance for both. We have implemented our system using Xen and Hadoop, and our evaluation shows that our schedulers allow a mixed cluster to reduce web response times by more than ten fold, while meeting more Hadoop deadlines and lowering total task execution times by 6.5%.
Wei Zhang 0052, Sundaresan Rajasekaran, Timothy Wood 0001, Mingfa Zhu
CCGRID3
2014 Topology Discovery and Service Classification for Distributed-Aware Clouds
abstract
Cloud data centers are difficult to manage because providers have no knowledge of what applications are being run by customers or how they interact. As a consequence, current clouds provide minimal automated management functionality, passing the problem on to users who have access to even fewer tools since they lack insight into the underlying infrastructure. Ideally, the cloud platform, not the customer, should be managing data center resources in order to both use them efficiently and provide strong application-level performance and reliability guarantees. To do this, we believe that clouds must become "distibuted-aware" so that they can deduce the overall structure and dependencies within a client's distributed applications and use that knowledge to better guide management services. Towards this end we are developing a light-weight topology detection system that maps distributed applications and a service classification algorithm that can determine not only overall application types, but individual VM roles as well.
Jinho Hwang, Guyue Liu, Sai Zeng, Frederick Y. Wu, Timothy Wood 0001
IC2E5
2014 NetVM: High Performance and Flexible Networking Using Virtualization on Commodity Platforms
Jinho Hwang, K. K. Ramakrishnan, Timothy Wood 0001
NSDI3
2014 Mortar: filling the gaps in data center memory
abstract
Data center servers are typically overprovisioned, leaving spare memory and CPU capacity idle to handle unpredictable workload bursts by the virtual machines running on them. While this allows for fast hotspot mitigation, it is also wasteful. Unfortunately, making use of spare capacity without impacting active applications is particularly difficult for memory since it typically must be allocated in coarse chunks over long timescales. In this work we propose re- purposing the poorly utilized memory in a data center to store a volatile data store that is managed by the hypervisor. We present two uses for our Mortar framework: as a cache for prefetching disk blocks, and as an application-level distributed cache that follows the memcached protocol. Both prototypes use the framework to ask the hypervisor to store useful, but recoverable data within its free memory pool. This allows the hypervisor to control eviction policies and prioritize access to the cache. We demonstrate the benefits of our prototypes using realistic web applications and disk benchmarks, as well as memory traces gathered from live servers in our university's IT department. By expanding and contracting the data store size based on the free memory available, Mortar improves average response time of a web application by up to 35% compared to a fixed size memcached deployment, and improves overall video streaming performance by 45% through prefetching.
Jinho Hwang, Ahsen J. Uppal, Timothy Wood 0001, H. Howie Huang
VEE3
2014 Cost-Aware Cloud Bursting for Enterprise Applications
abstract
The high cost of provisioning resources to meet peak application demands has led to the widespread adoption of pay-as-you-go cloud computing services to handle workload fluctuations. Some enterprises with existing IT infrastructure employ a hybrid cloud model where the enterprise uses its own private resources for the majority of its computing, but then “bursts” into the cloud when local resources are insufficient. However, current commercial tools rely heavily on the system administrator’s knowledge to answer key questions such as when a cloud burst is needed and which applications must be moved to the cloud. In this article, we describe Seagull, a system designed to facilitate cloud bursting by determining which applications should be transitioned into the cloud and automating the movement process at the proper time. Seagull optimizes the bursting of applications using an optimization algorithm as well as a more efficient but approximate greedy heuristic. Seagull also optimizes the overhead of deploying applications into the cloud using an intelligent precopying mechanism that proactively replicates virtualized applications, lowering the bursting time from hours to minutes. Our evaluation shows over 100% improvement compared to naïve solutions but produces more expensive solutions compared to ILP. However, the scalability of our greedy algorithm is dramatically better as the number of VMs increase. Our evaluation illustrates scenarios where our prototype can reduce cloud costs by more than 45% when bursting to the cloud, and that the incremental cost added by precopying applications is offset by a burst time reduction of nearly 95%.
Tian Guo 0001, Upendra Sharma, Prashant J. Shenoy, Timothy Wood 0001, Sambit Sahu
ACM Trans. Internet Techn.4
2013 Mortar: filling the gaps in data center memory
abstract
Data center servers are typically overprovisioned, leaving spare memory and CPU capacity idle to handle unpredictable workload bursts by the virtual machines running on them [1, 2, 3]. While this allows for fast hotspot mitigation, it is also wasteful. Unfortunately, making use of spare capacity without impacting active applications is particularly difficult for memory since it typically must be allocated in coarse chunks over long timescales [4, 5, 6, 7]. In this work we propose repurposing the poorly utilized memory in a data center to store a volatile data store that is managed by the hypervisor. We present two uses for our Mortar framework: as a cache for prefetching disk blocks [8, 9, 10], and as an application-level distributed cache that follows the memcached protocol [11, 12]. Both prototypes use the framework to ask the hypervisor to store useful, but recoverable data within its free memory pool. This allows the hypervisor to control eviction policies and prioritize access to the cache.
Jinho Hwang, Ahsen J. Uppal, Timothy Wood 0001, H. Howie Huang
SoCC3
2013 HybridMR: A Hierarchical MapReduce Scheduler for Hybrid Data Centers
abstract
Virtualized environments are attractive because they simplify cluster management, while facilitating cost-effective workload consolidation. As a result, virtual machines in public clouds or private data centers, have become the norm for running transactional applications like web services and virtual desktops. On the other hand, batch workloads like MapReduce, are typically deployed in a native cluster to avoid the performance overheads of virtualization. While both these virtual and native environments have their own strengths and weaknesses, we demonstrate in this work that it is feasible to provide the best of these two computing paradigms in a hybrid platform. In this paper, we make a case for a hybrid data center consisting of native and virtual environments, and propose a 2-phase hierarchical scheduler, called HybridMR, for the effective resource management of interactive and batch workloads. In the first phase, HybridMR classifies incoming MapReduce jobs based on the expected virtualization overheads, and uses this information to automatically guide placement between physical and virtual machines. In the second phase, HybridMR manages the run-time performance of MapReduce jobs collocated with interactive applications in order to provide best effort delivery to batch jobs, while complying with the Service Level Agreements (SLAs) of interactive applications. By consolidating batch jobs with over-provisioned foreground applications, the available unused resources are better utilized, resulting in improved application performance and energy efficiency. Evaluations on a hybrid cluster consisting of 24 physical servers and 48 virtual machines, with diverse workload mix of interactive and batch MapReduce applications, demonstrate that HybridMR can achieve up to 40% improvement in the completion times of MapReduce jobs, over the virtual-only case, while complying with the SLAs of interactive applications. Compared to the native-only cluster, at the cost of minimal performance penalty, HybridMR boosts resource utilization by 45%, and achieves up to 43% energy savings. These results indicate that a hybrid data center with an efficient scheduling mechanism can provide a cost-effective solution for hosting both batch and interactive workloads.
Bikash Sharma, Timothy Wood 0001, Chita R. Das
ICDCS2
2013 A component-based performance comparison of four hypervisors
Jinho Hwang, Sai Zeng, Frederick Wu, Timothy Wood 0001
IM4
2013 Benefits and challenges of managing heterogeneous data centers
Jinho Hwang, Sai Zeng, Frederick Wu, Timothy Wood 0001
IM4
2013 Firewall performance optimization using data mining techniques
abstract
This paper presents a novel approach to improve firewall performance using data mining techniques. A traditional packet filtering firewall compares a packet against each filtering rule until a match is found. The filtering rules are stored as a rule list. Therefore, the time required to process a packet depends linearly on the number of filtering rules. This time can be prohibitively large for a firewall containing hundreds of rules and the firewall can be a bottleneck for the network if high bandwidth is required. To enhance the firewall performance, we propose a data mining solution. In this approach, instead of comparing the packet with each of the filtering rules, the firewall predicts which rule is most likely going to match the packet. This significantly reduces the processing time taken by the firewall to filter each packet and thus improves its performance. Comparisons were made between the cumulative processing time taken by a standard firewall and the enhanced firewall with data mining to process millions of packets. Compared to the standard firewall, the enhanced firewall took 40% less time in processing the packets.
Umniya Mustafa, Mohammad M. Masud 0001, Zouheir Trabelsi, Timothy Wood 0001, Zainab Al Harthi
IWCMC4
2012 Adaptive dynamic priority scheduling for virtual desktop infrastructures
abstract
Virtual Desktop Infrastructures (VDIs) are gaining popularity in cloud computing by allowing companies to deploy their office environments in a virtualized setting instead of relying on physical desktop machines. Consolidating many users into a VDI environment can significantly lower IT management expenses and enables new features such as “available-anywhere” desktops. However, barriers to broad adoption include the slow performance of virtualized I/O, CPU scheduling interference problems, and shared-cache contention. In this paper, we propose a new soft real-time scheduling algorithm that employs flexible priority designations (via utility functions) and automated scheduler class detection (via hypervisor monitoring of user behavior) to provide a higher quality user experience. We have implemented our scheduler within the Xen virtualization platform, and demonstrate that the overheads incurred from co-locating large numbers of virtual machines can be reduced from 66% with existing schedulers to under 2% in our system. We evaluate the benefits and overheads of using a smaller scheduling time quantum in a VDI setting, and show that the average overhead time per scheduler call is on the same order as the existing SEDF and Credit schedulers.
Jinho Hwang, Timothy Wood 0001
IWQoS2
2012 An Empirical Study of Memory Sharing in Virtual Machines
Sean Kenneth Barker, Timothy Wood 0001, Prashant J. Shenoy, Ramesh K. Sitaraman
USENIX ATC2
2012 Seagull: Intelligent Cloud Bursting for Enterprise Applications
Tian Guo 0001, Upendra Sharma, Timothy Wood 0001, Sambit Sahu, Prashant J. Shenoy
USENIX ATC3
2012 Enterprise-Ready Virtual Cloud Pools: Vision, Opportunities and Challenges
abstract
Cloud computing platforms such as Amazon EC2 provide customers with flexible, on demand resources at low cost. However, while existing offerings are useful for providing basic computation and storage resources, they have not provided the transparency, security and network controls that many enterpise customers would like. While cloud computing has a great potential to change how enterprises run and manage their IT systems, a more comprehensive control over network resources and security needs to be provided for such users. Towards this goal, we propose a Virtual Cloud Pool abstraction to logically unify cloud and enterprise data center resources, and present the vision behind CloudNet, a cloud platform architecture which utilizes virtual private networks to securely and seamlessly link cloud and enterprise sites. It also enables the pooling of resources across data centers to provide enterprises the capability to have cloud resources that are dynamic and adaptive to their needs. We describe several usage scenarios for virtual cloud pools and discuss the benefits of using this abstraction in enterprise settings.
Timothy Wood 0001, K. K. Ramakrishnan, Prashant J. Shenoy, Jacobus E. van der Merwe
Comput. J.1
2012 Modellus: Automated modeling of complex internet data center applications
abstract
The rising complexity of distributed server applications in Internet data centers has made the tasks of modeling and analyzing their behavior increasingly difficult. This article presents Modellus , a novel system for automated modeling of complex web-based data center applications using methods from queuing theory, data mining, and machine learning. Modellus uses queuing theory and statistical methods to automatically derive models to predict the resource usage of an application and the workload it triggers; these models can be composed to capture multiple dependencies between interacting applications. Model accuracy is maintained by fast, distributed testing, automated relearning of models when they change, and methods to bound prediction errors in composite models. We have implemented a prototype of Modellus, deployed it on a data center testbed, and evaluated its efficacy for modeling and analysis of several distributed multitier web applications. Our results show that this feature-based modeling technique is able to make predictions across several data center tiers, and maintain predictive accuracy (typically 95% or better) in the face of significant shifts in workload composition; we also demonstrate practical applications of the Modellus system to prediction and provisioning of real-world data center applications.
Peter Desnoyers, Timothy Wood 0001, Prashant J. Shenoy, Sangameshwar Patil, Harrick M. Vin
ACM Trans. Web2
2011 PipeCloud: using causality to overcome speed-of-light delays in cloud-based disaster recovery
abstract
Disaster Recovery (DR) is a desirable feature for all enterprises, and a crucial one for many. However, adoption of DR remains limited due to the stark tradeoffs it imposes. To recover an application to the point of crash, one is limited by financial considerations, substantial application overhead, or minimal geographical separation between the primary and recovery sites. In this paper, we argue for cloud-based DR and pipelined synchronous replication as an antidote to these problems. Cloud hosting promises economies of scale and on-demand provisioning that are a perfect fit for the infrequent yet urgent needs of DR. Pipelined synchrony addresses the impact of WAN replication latency on performance, by efficiently overlapping replication with application processing for multi-tier servers. By tracking the consequences of the disk modifications that are persisted to a recovery site all the way to client-directed messages, applications realize forward progress while retaining full consistency guarantees for client-visible state in the event of a disaster. PipeCloud, our prototype, is able to sustain these guarantees for multi-node servers composed of black-box VMs, with no need of application modification, resulting in a perfect fit for the arbitrary nature of VM-based cloud hosting. We demonstrate disaster failover to the Amazon EC2 platform, and show that PipeCloud can increase throughput by an order of magnitude and reduce response times by more than half compared to synchronous replication, all while providing the same zero data loss consistency guarantees.
Timothy Wood 0001, H. Andrés Lagar-Cavilla, K. K. Ramakrishnan, Prashant J. Shenoy, Jacobus E. van der Merwe
SoCC1
2011 ZZ and the art of practical BFT execution
abstract
The high replication cost of Byzantine fault-tolerance (BFT) methods has been a major barrier to their widespread adop-tion in commercial distributed applications. We present ZZ, a new approach that reduces the replication cost of BFT ser-vices from 2f + 1 to practically f + 1. The key insight in ZZ is to use f + 1 execution replicas in the normal case and to activate additional replicas only upon failures. In data cen-ters where multiple applications share a physical server, ZZ reduces the aggregate number of execution replicas running in the data center, improving throughput and response times. ZZ relies on virtualization—a technology already employed in modern data centers—for fast replica activation upon fail-ures, and enables newly activated replicas to immediately be-gin processing requests by fetching state on-demand. A pro-totype implementation of ZZ using the BASE library and Xen shows that, when compared to a system with 2f + 1 repli-cas, our approach yields lower response times and up to 33% higher throughput in a prototype data center with four BFT web applications. We also show that ZZ can handle simulta-neous failures and achieve sub-second recovery. 1
Timothy Wood 0001, Arun Venkataramani, Prashant J. Shenoy, Emmanuel Cecchet
EuroSys1
2011 CloudNet: dynamic pooling of cloud resources by live WAN migration of virtual machines
abstract
Virtual machine technology and the ease with which VMs can be migrated within the LAN, has changed the scope of resource management from allocating resources on a single server to manipulating pools of resources within a data center. We expect WAN migration of virtual machines to likewise transform the scope of provisioning compute resources from a single data center to multiple data centers spread across the country or around the world. In this paper we present the CloudNet architecure as a cloud framework consisting of cloud computing platforms linked with a VPN based network infrastructure to provide seamless and secure connectivity between enterprise and cloud data center sites. To realize our vision of efficiently pooling geographically distributed data center resources, CloudNet provides optimized support for live WAN migration of virtual machines. Specifically, we present a set of optimizations that minimize the cost of transferring storage and virtual machine memory during migrations over low bandwidth and high latency Internet links. We evaluate our system on an operational cloud platform distributed across the continental US. During simultaneous migrations of four VMs between data centers in Texas and Illinois, CloudNet's optimizations reduce memory migration time by 65% and lower bandwidth consumption for the storage and memory transfer by 19GB, a 50% reduction.
Timothy Wood 0001, K. K. Ramakrishnan, Prashant J. Shenoy, Jacobus E. van der Merwe
VEE1
2009 Memory buddies: exploiting page sharing for smart colocation in virtualized data centers
abstract
Many data center virtualization solutions, such as VMware ESX, employ content-based page sharing to consolidate the resources of multiple servers. Page sharing identifies virtual machine memory pages with identical content and consolidates them into a single shared page. This technique, implemented at the host level, applies only between VMs placed on a given physical host. In a multi-server data center, opportunities for sharing may be lost because the VMs holding identical pages are resident on different hosts. In order to obtain the full benefit of content-based page sharing it is necessary to place virtual machines such that VMs with similar memory content are located on the same hosts.
Timothy Wood 0001, Gabriel Tarasuk-Levin, Prashant J. Shenoy, Peter Desnoyers, Emmanuel Cecchet, Mark D. Corner
VEE1
2009 Sandpiper: Black-box and gray-box resource management for virtual machines
Timothy Wood 0001, Prashant J. Shenoy, Arun Venkataramani, Mazin S. Yousif
Comput. Networks1
2008 Profiling and Modeling Resource Usage of Virtualized Applications
Timothy Wood 0001, Ludmila Cherkasova, Kivanc M. Ozonat, Prashant J. Shenoy
Middleware1
2008 Agile dynamic provisioning of multi-tier Internet applications
abstract
Dynamic capacity provisioning is a useful technique for handling the multi-time-scale variations seen in Internet workloads. In this article, we propose a novel dynamic provisioning technique for multi-tier Internet applications that employs (1) a flexible queuing model to determine how much of the resources to allocate to each tier of the application, and (2) a combination of predictive and reactive methods that determine when to provision these resources, both at large and small time scales. We propose a novel data center architecture based on virtual machine monitors to reduce provisioning overheads. Our experiments on a forty-machine Xen/Linux-based hosting platform demonstrate the responsiveness of our technique in handling dynamic workloads. In one scenario where a flash crowd caused the workload of a three-tier application to double, our technique was able to double the application capacity within five minutes, thus maintaining response-time targets. Our technique also reduced the overhead of switching servers across applications from several minutes to less than a second, while meeting the performance targets of residual sessions.
Bhuvan Urgaonkar, Prashant J. Shenoy, Abhishek Chandra, Pawan Goyal 0001, Timothy Wood 0001
ACM Trans. Auton. Adapt. Syst.5
2007 Black-box and Gray-box Strategies for Virtual Machine Migration
Timothy Wood 0001, Prashant J. Shenoy, Arun Venkataramani, Mazin S. Yousif
NSDI1
2005 The feasibility of launching and detecting jamming attacks in wireless networks
abstract
Wireless networks are built upon a shared medium that makes it easy for adversaries to launch jamming-style attacks. These attacks can be easily accomplished by an adversary emitting radio frequency signals that do not follow an underlying MAC protocol. Jamming attacks can severely interfere with the normal operation of wireless networks and, consequently, mechanisms are needed that can cope with jamming attacks. In this paper, we examine radio interference attacks from both sides of the issue: first, we study the problem of conducting radio interference attacks on wireless networks, and second we examine the critical issue of diagnosing the presence of jamming attacks. Specifically, we propose four different jamming attack models that can be used by an adversary to disable the operation of a wireless network, and evaluate their effectiveness in terms of how each method affects the ability of a wireless node to send and receive packets. We then discuss different measurements that serve as the basis for detecting a jamming attack, and explore scenarios where each measurement by itself is not enough to reliably classify the presence of a jamming attack. In particular, we observe that signal strength and carrier sensing time are unable to conclusively detect the presence of a jammer. Further, we observe that although by using packet delivery ratio we may differentiate between congested and jammed scenarios, we are nonetheless unable to conclude whether poor link utility is due to jamming or the mobility of nodes. The fact that no single measurement is sufficient for reliably classifying the presence of a jammer is an important observation, and necessitates the development of enhanced detection schemes that can remove ambiguity when detecting a jammer. To address this need, we propose two enhanced detection protocols that employ consistency checking. The first scheme employs signal strength measurements as a reactive consistency check for poor packet delivery ratios, while the second scheme employs location information to serve as the consistency check. Throughout our discussions, we examine the feasibility and effectiveness of jamming attacks and detection schemes using the MICA2 Mote platform.
Wenyuan Xu 0001, Wade Trappe, Yanyong Zhang, Timothy Wood 0001
MobiHoc4