Lin Wang 0015

dblp:17/6729-15 · DBLP profile ↗
← Back
71ranked-venue papers
13as first author
30since 2021 · last 2026
0000-0001-7181-6128ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 36 · 7 first-author · 15 since 2021Systems, architecture and hardware · 22 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Artificial intelligence and machine learning · 3Theory of computation · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Cells on Autopilot: Adaptive Cell (Re)Selection via Reinforcement Learning
Marvin Illian, Ramin Khalili, Antônio Augusto de Aragão Rocha, Lin Wang 0015
WiOpt4
2026 FreeBeacon: Efficient Communication and Data Aggregation in Battery-Free IoT
Gaosheng Liu, Kasim Sinan Yildirim, Lin Wang 0015
IEEE Trans. Mob. Comput.3
2025 It Takes Two to Tango: Serverless Workflow Serving via Bilaterally Engaged Resource Adaptation
abstract
Serverless platforms typically adopt an earlybinding approach for function sizing, requiring developers to specify an immutable size for each function within a workflow beforehand. Accounting for potential runtime variability, developers must size functions for worst-case scenarios to ensure service-level objectives (SLOs), resulting in significant resource inefficiency. To address this issue, we propose Janus, a novel resource adaptation framework for serverless platforms. Janus employs a late-binding approach, allowing function sizes to be dynamically adapted based on runtime conditions. The main challenge lies in the information barrier between the developer and the provider: developers lack access to runtime information, while providers lack domain knowledge about the workflow. To bridge this gap, Janus allows developers to provide hints containing rules and options for resource adaptation. Providers then follow these hints to dynamically adjust resource allocation at runtime based on real-time function execution information, ensuring compliance with SLOs. We implement Janus and conduct extensive experiments with real-world serverless workflows. Our results demonstrate that Janus enhances resource efficiency by up to 34.7% compared to the state-of-the-art.
Jing Wu 0024, Lin Wang 0015, Quanfeng Deng, Chen Yu 0003, Bingheng Yan, Fangming Liu
IPDPS2
2025 Uirapuru: Timely Video Analytics for High-Resolution Steerable Cameras on Edge Devices
abstract
Real-time video analytics on high-resolution cameras has become a popular technology for various intelligent services like traffic control and crowd monitoring. While extensive work has been done on improving analytics accuracy with timing guarantees, virtually all of them target static viewpoint cameras. In this paper, we present Uirapuru, a novel framework for real-time, edge-based video analytics on high-resolution steerable cameras. The actuation performed by those cameras brings significant dynamism to the scene, presenting a critical challenge to existing popular approaches such as frame tiling. To address this problem, Uirapuru incorporates a comprehensive understanding of camera actuation into the system design paired with fast adaptive tiling at a per-frame level. We evaluate Uirapuru on a high-resolution video dataset, augmented by pan-tilt-zoom (PTZ) movements typical for steerable cameras and on real-world videos collected from an actual PTZ camera. Our experimental results show that Uirapuru provides up to 1.45× improvement in accuracy while respecting specified latency budgets or reaches up to 4.53× inference speedup with on-par accuracy compared to state-of-the-art static camera approaches.
Guilherme Henrique Apostolo, Pablo Bauszat, Vinod Nigade, Henri E. Bal, Lin Wang 0015
MobiCom5
2025 OC-HMAS: Dynamic Self-Organization and Self-Correction in Heterogeneous Multiagent Systems Using Multimodal Large Models
abstract
Heterogeneous multiagent systems (HMASs) leverage diverse agent capabilities to address complex tasks in dynamic environments, yet traditional approaches face limitations in autonomy and generalization when adapting to evolving scenarios. To overcome these challenges, we propose OC-HMAS, an IoT-integrated framework that synergizes self-organization and self-correction through multimodal perception. The system processes RGB images, LiDAR point clouds, and instance segmentation maps for real-time environmental awareness, while vision-language models and large language models (LLMs) jointly enable context-aware task decomposition, role allocation, and adaptive planning. Integrated path optimization and obstacle avoidance mechanisms further ensure operational safety and scalability across logistics, inspection, and search-and-rescue operations. Experimental validation demonstrates the framework’s superiority over SMRC-LLM, with 5.15% higher success rates and 14.2% faster task completion in logistics, alongside 4.69% accuracy gains and 12.9% time reduction in inspection scenarios. These results validate its enhanced adaptability in IoT-augmented environments, establishing a new benchmark for autonomous HMAS deployment.
Ping Feng, Tingting Yang 0001, Mingyang Liang, Lin Wang 0015, Yuan Gao 0024
IEEE Internet Things J.4
2025 Working Smarter Not Harder: Hybrid Cooling for Deep Learning in Edge Datacenters
abstract
The proliferation of deep-learning-based mobile and IoT applications has driven the increasing deployment of edge datacenters equipped with domain-specific accelerators. The unprecedented computing power offered by these accelerators puts a heavy burden on the cooling system, motivating more potent cooling techniques like cold water cooling. However, we observe that cold water cooling results in significant energy waste in edge datacenters due to the fluctuating resource utilization both spatially and temporally. To tackle this issue, we propose the concept of “working smarter” by slowing down accelerators deliberately whenever possible and enabling warm water cooling during these times to achieve cooling efficiency. Based on this concept, we develop Hyco—a hybrid water cooling system tailored for edge datacenters running deep learning workloads. First, Hyco features a zone-based cooling architecture enabling dynamic switching between cold water and warm water cooling. Then, based on a lightweight latency estimation method, Hyco incorporates a learning-based scheduling scheme to determine “which” accelerator workers and “when” to slow down through an adaptive and intelligent power-latency trade-off for deep learning models. The simulation with real-world traces shows that Hyco reduces the cooling energy consumption by up to 34.74× while satisfying latency constraints more than 99% of the time for deep-learning-based applications.
Qiangyu Pei, Yongjie Yuan, Haichuan Hu, Lin Wang 0015, Bingheng Yan, Chen Yu 0003, Fangming Liu
IEEE Trans. Sustain. Comput.4
2024 A Little Certainty is All We Need: Discovery and Synchronization Acceleration in Battery-Free IoT
abstract
The vision of sustainable IoT constructed from battery-free devices has attracted ample interest in the research community. Yet, efficient device discovery and synchronization—a fundamental problem in IoT systems—remains a critical challenge mainly due to the uncertain ambient energy availability across battery-free devices. We argue that bringing in a small level of certainty is necessary for facilitating communication in battery-free IoT. We propose Pulsar where we introduce a small number of battery-powered devices, serving as the communication coordinator for a large number of battery-free devices. We develop two communication schemes, namely one-to-one, and all-to-all, for Pulsar. Our results based on simulations and prototype-based experiments show that Pulsar achieves consistently good performance across different scenarios while requiring no special hardware or environmental conditions.
Gaosheng Liu, Vinod Nigade, Henri E. Bal, Lin Wang 0015
APNet4
2024 InferCool: Enhancing AI Inference Cooling through Transparent, Non-Intrusive Task Reassignment
abstract
The increasing power consumption of AI inference in modern datacenters has escalated cooling demands significantly, necessitating the adoption of potent cooling approaches like water cooling. Unlike traditional cloud workloads, AI inference has unique characteristics that create substantial gaps in achieving optimal cooling efficiency. In this work, we present the first comprehensive measurement study of AI inference cooling across various models within an industrial-ready scheduling framework, highlighting significant inefficiencies and their causes. To fill the gap while following the fundamental requirements of cooling systems, we explore a new opportunity presented by modern Multi-Instance GPU-enabled inference serving, where the scheduling dimension is naturally orthogonal to the cooling dimension. Building on this insight, we develop InferCool, a cooling middleware designed to enhance cooling efficiency for inference serving through transparent, non-intrusive task reassignment. It includes a streamlined power and temperature prediction approach and a thermal-aware, adaptive application deployment and request scheduling mechanism. Real-world experiments on a water-cooled testbed and a three-node cluster demonstrate that InferCool can reduce the maximum GPU temperature by 5°C across eight A100 GPUs, equivalent to cooling energy savings of about 20%. Importantly, InferCool requires no modifications to existing cooling infrastructures and is compatible with existing scheduling systems.
Qiangyu Pei, Lin Wang 0015, Bingheng Yan, Chen Yu 0003, Fangming Liu
SoCC2
2024 Train Once Apply Anywhere: Effective Scheduling for Network Function Chains Running on FUMES
abstract
The emergence of network function virtualization has enabled network function chaining as a flexible approach for building complex network services. However, the high degree of flexibility envisioned for orchestrating network function chains introduces several challenges to support dynamism in workloads and the environment necessary for their realization. Existing works mostly consider supporting dynamism by re-adjusting provisioning of network function instances, incurring reaction times that are prohibitively high in practice. Existing solutions to dynamic packet scheduling rely on centralized schedulers and a priori knowledge of traffic characteristics, and cannot handle changes in the environment like link failures.We fill this gap by presenting FUMES, a reinforcement learning based distributed agent design for the runtime scheduling problem of assigning packets undergoing treatment by network function chains to network function instances. Our design consists of multiple distributed agents that cooperatively work on the scheduling problem. A key design choice enables agents, once trained, to be applicable for unknown chains and traffic patterns including branching, and different environments including link failures. The paper presents the system design and shows its suitability for realistic deployments. We empirically compare FUMES with state-of-the-art runtime scheduling solutions showing improved scheduling decisions at lower server capacity.
Marcel Blöcher, Nils Nedderhut, Pavel Chuprikov, Ramin Khalili, Patrick Eugster, Lin Wang 0015
INFOCOM6
2024 X-Stream: A Flexible, Adaptive Video Transformer for Privacy-Preserving Video Stream Analytics
abstract
Video stream analytics (VSA) systems fuel many exciting applications that facilitate people’s lives, but also raise critical concerns about exposing too much individuals’ privacy. To alleviate these concerns, various frameworks have been presented to enhance the privacy of VSA systems. Yet, existing solutions suffer two limitations: (1) being scenario-customized, thus limiting the generality of adapting to multifarious scenarios, (2) requiring complex, imperative programming, and tedious process, thus largely reducing the usability of such systems. In this paper, we present X-Stream, a privacy-preserving video transformer that achieves flexibility and efficiency for a large variety of VSA tasks. X-Stream features three major novel designs: (1) a declarative query interface that provides a simple yet expressive interface for users to describe both their privacy protection and content exposure requirements, (2) an adaptation mechanism that dynamically selects the most suitable privacy-preserving techniques and their parameters based on the current video context, and (3) an efficient execution engine that incorporates optimizations for multi-task deduplication and inter-frame inference. We implement X-Stream and evaluate it with representative VSA tasks and public video datasets. The results show that X-Stream achieves significantly improved privacy protection quality and performance over the state-of-the-art, while being simple to use.
Dou Feng, Lin Wang 0015, Lingching Tung, Fangming Liu
INFOCOM2
2024 NetNN: Neural Intrusion Detection System in Programmable Networks
abstract
The rise of deep learning has led to various successful attempts to apply deep neural networks (DNNs) for important networking tasks such as intrusion detection. Yet, running DNNs in the network control plane, as typically done in existing proposals, suffers from high latency that impedes the practicality of such approaches. This paper introduces NetNN, a novel DNN-based intrusion detection system that runs completely in the network data plane to achieve low latency. NetNN adopts raw packet information as input, avoiding complicated feature engineering. NetNN mimics the DNN dataflow execution by mapping DNN parts to a network of programmable switches, executing partial DNN computations on individual switches, and generating packets carrying intermediate execution results between these switches. We implement NetNN in P4 and demonstrate the feasibility of such an approach. Experimental results show that NetNN can improve the intrusion detection accuracy to 99% while meeting the real-time requirement.
Kamran Razavi, Shayan Davari Fard, George Karlos, Vinod Nigade, Max Mühlhäuser, Lin Wang 0015
ISCC6
2024 NetCL: A Unified Programming Framework for In-Network Computing
abstract
The emergence of programmable data planes (PDPs) has paved the way for in-network computing (INC), a paradigm wherein networking devices actively participate in distributed computations. However, PDPs are still a niche technology, mostly available to network operators, and rely on packet-processing DSLs like P4. This necessitates great networking expertise from INC programmers to articulate computational tasks in networking terms and reason about their code. To lift this barrier to INC we propose a unified compute interface for the data plane. We introduce $\mathrm{C} / \mathrm{C}++$ extensions that allow INC to be expressed as kernel functions processing in-flight messages, and APIs for establishing INC-aware communication. We develop a compiler that translates kernels into P4, and thin runtimes that handle the required network plumbing, shielding INC programmers from low-level networking details. We evaluate our system using common INC applications from the literature.
George Karlos, Henri E. Bal, Lin Wang 0015
SC3
2024 λGrapher: A Resource-Efficient Serverless System for GNN Serving through Graph Sharing
abstract
Graph Neural Networks (GNNs) have been increasingly adopted for graph analysis in web applications such as social networks. Yet, efficient GNN serving remains a critical challenge due to high workload fluctuations and intricate GNN operations. Serverless computing, thanks to its flexibility and agility, offers on-demand serving of GNN inference requests. Alas, the request-centric serverless model is still too coarse-grained to avoid resource waste.
Haichuan Hu, Fangming Liu, Qiangyu Pei, Yongjie Yuan, Zichen Xu 0001, Lin Wang 0015
WWW6
2024 Inference serving with end-to-end latency SLOs over dynamic edge networks
abstract
Abstract While high accuracy is of paramount importance for deep learning (DL) inference, serving inference requests on time is equally critical but has not been carefully studied especially when the request has to be served over a dynamic wireless network at the edge. In this paper, we propose Jellyfish—a novel edge DL inference serving system that achieves soft guarantees for end-to-end inference latency service-level objectives (SLO). Jellyfish handles the network variability by utilizing both data and deep neural network (DNN) adaptation to conduct tradeoffs between accuracy and latency. Jellyfish features a new design that enables collective adaptation policies where the decisions for data and DNN adaptations are aligned and coordinated among multiple users with varying network conditions. We propose efficient algorithms to continuously map users and adapt DNNs at runtime, so that we fulfill latency SLOs while maximizing the overall inference accuracy. We further investigate dynamic DNNs, i.e., DNNs that encompass multiple architecture variants, and demonstrate their potential benefit through preliminary experiments. Our experiments based on a prototype implementation and real-world WiFi and LTE network traces show that Jellyfish can meet latency SLOs at around the 99th percentile while maintaining high accuracy.
Vinod Nigade, Pablo Bauszat, Henri E. Bal, Lin Wang 0015
Real Time Syst.4
2024 Data on the Go: Seamless Data Routing for Intermittently-Powered Battery-Free Sensing
abstract
The rising demand for sustainable IoT has promoted the adoption of battery-free devices intermittently powered by ambient energy for sensing. However, the intermittency poses significant challenges in sensing data collection. Despite recent efforts to enable one-to-one communication, routing data across multiple intermittently-powered battery-free devices, a crucial requirement for a sensing system, remains a formidable challenge. This paper fills this gap by introducing Swift, which enables seamless data routing in intermittently-powered battery-free sensing systems. Swift overcomes the challenges posed by device intermittency and heterogeneous energy conditions through three major innovative designs. First, Swift incorporates a reliable node synchronization protocol backed by number theory, ensuring successful synchronization regardless of energy conditions. Second, Swift adopts a low-latency message forwarding protocol, allowing continuous message forwarding without repeated synchronization. Finally, Swift features a simple yet effective mechanism for routing path construction, enabling nodes to obtain the optimal path to the sink node with minimum hops. We implement Swift and perform large-scale experiments representing diverse real-world scenarios. The results demonstrate that Swift achieves an order of magnitude reduction in end-to-end message delivery time compared with the state-of-the-art approaches for intermittently-powered battery-free sensing systems.
Gaosheng Liu, Lin Wang 0015
IEEE Trans. Mob. Comput.2
2024 Graft: Efficient Inference Serving for Hybrid Deep Learning With SLO Guarantees via DNN Re-Alignment
abstract
Deep neural networks (DNNs) have been widely adopted for various mobile inference tasks, yet their ever-increasing computational demands are hindering their deployment on resource-constrained mobile devices. Hybrid deep learning partitions a DNN into two parts and deploys them across the mobile device and a server, aiming to reduce inference latency or prolong battery life of mobile devices. However, such partitioning produces (non-uniform) DNN fragments which are hard to serve efficiently on the server. This article presents Graft—an efficient inference serving system for hybrid deep learning with latency service-level objective (SLO) guarantees. Our main insight is to mitigate the non-uniformity by a core concept called DNN re-alignment, allowing multiple heterogeneous DNN fragments to be restructured to share layers. To fully exploit the potential of DNN re-alignment, Graft employs fine-grained GPU resource sharing. Based on that, we propose efficient algorithms for merging, grouping, and re-aligning DNN fragments to maximize request batching opportunities, minimizing resource consumption while guaranteeing the inference latency SLO. We implement a Graft prototype and perform extensive experiments with five types of widely used DNNs and real-world network traces. Our results show that Graft improves resource efficiency by up to 70% compared with the state-of-the-art inference serving systems.
Jing Wu 0024, Lin Wang 0015, Qirui Jin, Fangming Liu
IEEE Trans. Parallel Distributed Syst.2
2023 Routing for Intermittently-Powered Sensing Systems
abstract
Recently, intermittent computing (IC) has received tremendous attention due to its high potential in perpetual sensing for Internet-of-Things (IoT). By harvesting ambient energy, battery-free devices can perform sensing intermittently without maintenance, thus significantly improving IoT sustainability. To build a practical intermittently-powered sensing system, efficient routing across battery-free devices for data delivery is essential. However, the intermittency of these devices brings new challenges, rendering existing routing protocols inapplicable.In this paper, we propose RICS, a new routing scheme tailored for intermittently-powered sensing systems. RICS features two major designs to combat the intermittency challenge, with the goal of achieving low-latency data delivery on a network built with battery-free devices. First, RICS incorporates a fast topology construction protocol for each IC node to establish a path towards the sink node with the least hop count. Second, RICS employs a low-latency message forwarding protocol, which incorporates an efficient synchronization mechanism and a novel technique called pendulum-sync to avoid time-consuming repeated node synchronization. Our evaluation based on an implementation in OMNeT ++ and comprehensive experiments with varying system settings shows that RICS can achieve orders of magnitude latency reduction in data delivery compared with the state-of-the-art.
Gaosheng Liu, Lin Wang 0015
IPCCC2
2022 Dependency-Aware Traffic Management for Configuring On-demand in Service Meshes
abstract
Service mesh is a promising micro-services architecture due to its excellent governance capabilities. Unlike traditional service invocation, configurations for governance need to be issued in the service mesh. However, we find that the control-plane traffic of governance is distributed in full by default, i.e., each service in the data plane receives all configurations. The vast majority of the configurations are redundant for a specific service. Hence, it is important and challenging to make the control plane aware of the calling relationships between services. In this paper, we propose a traffic management mechanism named DATM. Using this mechanism, the entire cluster can be dynamically controlled and services can be configured on demand. It is implemented through a dependency-aware controller and monitors. The controller first processes the information listened to by the monitors and then analyzes the connection between the metrics and the service requests through intelligent algorithms. Finally, the control traffic for regulating the control plane is generated. Our proposed mechanism is experimentally compared with the default strategy and existing work across a wide set of load scenarios in a testbed based on Istio service mesh and Kubernetes. Experimental results demonstrate that our mechanism can save the storage resources of a single agent by 40% to 60%, and the number of cluster updates can be greatly reduced. From the perspective of the whole cluster, the optimization results are even better.
Lin Wang 0015, Xin Li 0005, Ning Wang 0018, Hao Li 0030, Xiaolin Qin, Jie Wu 0001
ICPADS1
2022 Optimal Admission Control Mechanism Design for Time-Sensitive Services in Edge Computing
abstract
Edge computing is a promising solution for reducing service latency by provisioning time-sensitive services directly from the network edge. However, upon workload peaks at the resource-limited edge, an edge service has to queue service requests, incurring high waiting time. Such quality of service (QoS) degradation ruins the reputation and reduces the long-term revenue of the service provider.To address this issue, we propose an admission control mechanism for time-sensitive edge services. Specifically, we allow the service provider to offer admission advice to arriving requests regarding whether to join for service or balk to seek alternatives. Our goal is twofold: maximizing revenue of the service provider and ensuring QoS if the provided admission advice is followed. To this end, we propose a threshold structure that estimates the highest length of the request queue. Leveraging such a threshold structure, we propose O2A, a mechanism to balance the trade-off between increasing revenue from accepting more requests and guaranteeing QoS by advising requests to balk. Rigorous analysis shows that O2A achieves the goal and that the provided admission advice is optimal for end-users to follow. We further validate O2A through trace-driven simulations with both synthetic and real-world service request traces.
Lin Wang 0015, Fangming Liu
INFOCOM2
2022 Retention-Aware Container Caching for Serverless Edge Computing
abstract
Serverless edge computing adopts an event-based model where Internet-of-Things (IoT) services are executed in lightweight containers only when requested, leading to significantly improved edge resource utilization. Unfortunately, the startup latency of containers degrades the responsiveness of IoT services dramatically. Container caching, while masking this latency, requires retaining resources thus compromising resource efficiency. In this paper, we study the retention-aware container caching problem in serverless edge computing. We leverage the distributed and heterogeneous nature of edge platforms and propose to optimize container caching jointly with request distribution. We reveal step by step that this joint optimization problem can be mapped to the classic ski-rental problem. We first present an online competitive algorithm for a special case where request distribution and container caching are based on a set of carefully designed probability distribution functions. Based on this algorithm, we propose an online algorithm called O-RDC for the general case, which incorporates the resource capacity and network latency by opportunistically distributing requests. We conduct extensive experiments to examine the performance of the proposed algorithms with both synthetic and real-world serverless computing traces. Our results show that ORDC outperforms existing caching strategies of current serverless computing platforms by up to 94.5% in terms of the overall system cost.
Lin Wang 0015, Fangming Liu
INFOCOM2
2022 FA2: Fast, Accurate Autoscaling for Serving Deep Learning Inference with SLA Guarantees
abstract
Deep learning (DL) inference has become an essential building block in modern intelligent applications. Due to the high computational intensity of DL, it is critical to scale DL inference serving systems in response to fluctuating workloads to achieve resource efficiency. Meanwhile, intelligent applications often require strict service level agreements (SLAs), which need to be guaranteed when the system is scaled. The problem is complex and has been tackled only in simple scenarios so far.This paper describes FA2, a fast and accurate autoscaler concept for DL inference serving systems. In contrast to related works, FA2 adopts a general, contrived two-phase approach. Specifically, it starts by capturing the autoscaling challenges in a comprehensive graph-based model. Then, FA2 applies targeted graph transformation and makes autoscaling decisions with an efficient algorithm based on dynamic programming. We implemented FA2 and built and evaluated a prototype. Compared with state-of-the-art autoscaling solutions, our experiments showed FA2 to achieve significant resource reduction (19% under CPUs and 25% under GPUs, on average) in combination with low SLA violations (less than 1.5%). FA2 performed close to the theoretical optimum, matching exactly the optimal decisions (with the least required resources) in 96.8% of all the cases in our evaluation.
Kamran Razavi, Manisha Luthra, Boris Koldehofe, Max Mühlhäuser, Lin Wang 0015
RTAS5
2022 Jellyfish: Timely Inference Serving for Dynamic Edge Networks
abstract
While high accuracy is of paramount importance for deep learning (DL) inference, serving inference requests on time is equally critical but has not been carefully studied especially when the request has to be served over a dynamic wireless network at the edge. In this paper, we propose Jellyfish—a novel edge DL inference serving system that achieves soft guarantees on end-to-end inference latency often specified as a service-level objective (SLO). To handle the network variability, Jellyfish exploits both data and deep neural network (DNN) adaptation to conduct tradeoffs between accuracy and latency. Jellyfish features a new design that enables collective adaptation policies where the decisions for data and DNN adaptations are aligned and coordinated among multiple users with varying network conditions. We propose efficient algorithms to dynamically adapt DNNs and map users, so that we fulfill latency SLOs while maximizing the overall inference accuracy. Our experiments based on a prototype implementation and real-world WiFi and LTE network traces show that Jellyfish can meet latency SLOs at around the 99th percentile while maintaining high accuracy.
Vinod Nigade, Pablo Bauszat, Henri E. Bal, Lin Wang 0015
RTSS4
2022 Healthor: Heterogeneity-aware Flow Control in DLTs to Increase Performance and Decentralization
abstract
Permissionless reputation-based distributed ledger technologies (DLTs) have been proposed to overcome blockchains’ shortcomings in terms of performance and scalability, and to enable feeless messages to power the machine-to-machine economy. These DLTs allow machines with widely heterogeneous capabilities to actively participate in message generation and consensus. However, the open nature of such DLTs can lead to the centralization of decision-making power, thus defeating the purpose of building a decentralized network. In this article, we introduce Healthor, a novel heterogeneity-aware flow-control mechanism for permissionless reputation-based DLTs. Healthor formalizes node heterogeneity by defining a health value as a function of its incoming message queue occupancy. We show that health signals can be used effectively by neighboring nodes to dynamically flow control messages while maintaining high decentralization. We perform extensive simulations, and show a 23% increase in throughput, a 76% decrease in latency and four times increased node participation in consensus compared to state-of-the-art. To the best of our knowledge, Healthor is the first system to systematically explore the ramifications of heterogeneity on DLTs and proposes a dynamic, heterogeneity-aware flow control. Healthor’s source code ( https://github.com/jonastheis/healthor ) and simulation result data set ( https://zenodo.org/record/4573698 ) are both publicly available.
Jonas Theis, Luigi Vigneri, Lin Wang 0015, Animesh Trivedi
Distributed Ledger Technol. Res. Pract.3
2022 Holistic Resource Scheduling for Data Center In-Network Computing
abstract
The recent trend towards more programmable switching hardware in data centers opens up new possibilities for distributed applications to leverage in-network computing (INC). Literature so far has largely focused on individual application scenarios of INC, leaving aside the problem of coordinating usage of potentially scarce and heterogeneous switch resources among multiple INC scenarios, applications, and users. Alas, the traditional model of resource pools of isolated compute containers does not fit an INC-enabled data center. This paper describes HIRE, a holistic INC-aware resource manager which allows for server-local and INC resources to be coordinated in unison. HIRE introduces a novel flexible resource (meta-)model to address heterogeneity and resource interchangeability, and includes two approaches for INC scheduling: (a) retrofitting existing schedulers; (b) designing a new one. For (a), HIRE presents a retrofitting API and demonstrates it with four state-of-the-art schedulers. For (b), HIRE proposes a flow-based scheduler, cast as a min-cost max-flow problem, where a unified cost model is used to integrate the different costs. Experiments with a workload trace of a 4000 machine cluster show that HIRE makes better use of INC resources by serving 8–30% more INC requests, while simultaneously reducing network detours by 20% and reducing tail placement latency by 50%.
Marcel Blöcher, Lin Wang 0015, Patrick Eugster, Max Schmidt
IEEE/ACM Trans. Netw.2
2022 EdgeDR: An Online Mechanism Design for Demand Response in Edge Clouds
abstract
The computing frontier is moving from centralized mega datacenters towards distributed cloudlets at the network edge. We argue that cloudlets are well-suited for handling power demand response to help the grid maintain stability due to more flexible workload management attributed to their distributed nature. However, they also require computing demand response to avoid overload and maintain reliability. To this end, we propose a novel online market mechanism, EdgeDR, to achieve cost efficiency in edge demand response programs. At a high level, we observe that the cloudlet operator can dynamically switch on/off entire cloudlets to compensate for the energy reduction required by the power grid or provide enough computing resources to the edge service. We formulate a long-term social cost minimization problem and decompose it into a series of one-round procurement auctions. In each auction instance, we propose to let the cloudlet tenants bid with cost functions of their two-dimension service quality degradation tolerance, and let the cloudlet operator choose the service quality, manage the workload, and schedule the cloudlet activation status. In addition, we present a dynamic payment mechanism for the operator to balance the tradeoff between short-term profit and long-term benefit in more practical scenarios. Via rigorous analysis, we exhibit that our bidding policy is individually rational and truthful; our workload management algorithm has near-optimal performance in each auction; and our overall online algorithm achieves a provable competitive ratio. We further confirm the performance of our mechanism through extensive trace-driven simulations.
Lei Jiao 0002, Fangming Liu, Lin Wang 0015
IEEE Trans. Parallel Distributed Syst.4
2022 HiTDL: High-Throughput Deep Learning Inference at the Hybrid Mobile Edge
abstract
Deep neural networks (DNNs) have become a critical component for inference in modern mobile applications, but the efficient provisioning of DNNs is non-trivial. Existing mobile- and server-based approaches compromise either the inference accuracy or latency. Instead, a hybrid approach can reap the benefits of the two by splitting the DNN at an appropriate layer and running the two parts separately on the mobile and the server respectively. Nevertheless, the DNN throughput in the hybrid approach has not been carefully examined, which is particularly important for edge servers where limited compute resources are shared among multiple DNNs. This article presents HiTDL, a runtime framework for managing multiple DNNs provisioned following the hybrid approach at the edge. HiTDL's mission is to improve edge resource efficiency by optimizing the combined throughput of all co-located DNNs, while still guaranteeing their SLAs. To this end, HiTDL first builds comprehensive performance models for DNN inference latency and throughout with respect to multiple factors including resource availability, DNN partition plan, and cross-DNN interference. HiTDL then uses these models to generate a set of candidate partition plans with SLA guarantees for each DNN. Finally, HiTDL makes global throughput-optimal resource allocation decisions by selecting partition plans from the candidate set for each DNN via solving a fairness-aware multiple-choice knapsack problem. Experimental results based on a prototype implementation show that HiTDL improves the overall throughput of the edge by$4.3\times$compared with the state-of-the-art.
Jing Wu 0024, Lin Wang 0015, Qiangyu Pei, Xingqi Cui, Fangming Liu, Tingting Yang 0001
IEEE Trans. Parallel Distributed Syst.2
2021 Switches for HIRE: resource scheduling for data center in-network computing
abstract
The recent trend towards more programmable switching hardware in data centers opens up new possibilities for distributed applications to leverage in-network computing (INC). Literature so far has largely focused on individual application scenarios of INC, leaving aside the problem of coordinating usage of potentially scarce and heterogeneous switch resources among multiple INC scenarios, applications, and users. The traditional model of resource pools of isolated compute containers does not fit an INC-enabled data center.
Marcel Blöcher, Lin Wang 0015, Patrick Eugster, Max Schmidt
ASPLOS2
2021 Don't You Worry 'Bout a Packet: Unified Programming for In-Network Computing
abstract
In-network computing is gaining momentum as programmable switches are increasingly employed for compute acceleration. Designed for packet processing, data plane programming languages force developers to express compute in networking terms, resulting in a complex, error-prone practice. We envision the unification of switch and host programming and propose the Net Compute Language (NCL), a C/C++ extension for expressing computational kernels for switches to execute. NCL implements Compute Centric Communication (C3), our proposed programming model for INC under which, point-to-point primitives are augmented to carry out computations. We motivate our approach with real-world use cases and discuss the technical challenges for its realization.
George Karlos, Henri E. Bal, Lin Wang 0015
HotNets3
2021 Better Never Than Late: Timely Edge Video Analytics Over the Air
abstract
Edge video analytics based on deep learning has become an important building block for many modern intelligent applications such as mobile augmented reality and autonomous driving. Various mechanisms have been developed to handle dynamic wireless networks, compute resource availability, and achieve high analytics accuracy via filtering, DNN compression, pruning, and adaptation. So far, limited attention has been paid to timeliness---providing strict service-level objectives (SLO) for edge video analytics pipelines, which is essential for the usability of user-interactive and mission-critical intelligent applications. In this paper, we analyze the challenges in achieving SLO for edge video analytics and present a system design for timely edge video analytics over the air leveraging a simple yet effective idea---feedback control. Our preliminary evaluation based on a system prototype and real-world network traces shows the potential of our design. We also discuss the limitations, calling for future work.
Vinod Nigade, Ramon Winder, Henri E. Bal, Lin Wang 0015
SenSys4
2021 Service Placement for Collaborative Edge Applications
abstract
Edge computing is emerging as a promising computing paradigm for supporting next-generation applications that rely on low-latency network connections in the Internet-of-Things (IoT) era. Many edge applications, such as multi-player augmented reality (AR) gaming and federated machine learning, require that distributed clients work collaboratively for a common goal through message exchanges. Given an edge network, it is an open problem how to deploy such collaborative edge applications to achieve the best overall system performance. This paper presents a formal study of this problem. We first provide a mix of cost models to capture the system. Based on a thorough formulation, we propose an iterative algorithm dubbed ITEM, where in each iteration, we construct a graph to encode all the costs and convert the cost optimization problem into a graph cut problem. By obtaining the minimum s-t cut via existing max-flow algorithms, we address the original problem via solving a series of graph cuts. We rigorously prove that ITEM has a parameterized constant approximation ratio. Inspired by the optimal stopping theory, we further design an online algorithm called OPTS, based on optimally alternating between partial and full placement updates. Our evaluations with real-world data traces demonstrate that ITEM performs close to the optimum (within 5%) and converges fast. OPTS achieves a bounded performance as expected while reducing full updates by more than 67% of the time.
Lin Wang 0015, Lei Jiao 0002, Ting He 0001, Jun Li 0001, Henri E. Bal
IEEE/ACM Trans. Netw.1
2020 Operator as a Service: Stateful Serverless Complex Event Processing
abstract
Complex Event Processing (CEP) is a powerful paradigm for scalable data management that is employed in many real-world scenarios such as detecting credit card fraud in banks. The so-called complex events are expressed using a specification language that is typically implemented and executed on a specific runtime system. While the tight coupling of these two components has been regarded as the key for supporting CEP at high performance, such dependencies pose several inherent challenges as follows. (1) Application development atop a CEP system requires extensive knowledge of how the runtime system operates, which is typically highly complex in nature. (2) The specification language dependence requires the need of domain experts and further restricts and steepens the learning curve for application developers.In this paper, we propose CEPless, a scalable data management system that decouples the specification from the runtime system by building on the principles of serverless computing. CEPless provides "operator as a service" and offers flexibility by enabling the development of CEP application in any specification language while abstracting away the complexity of the CEP runtime system. As part of CEPless, we designed and evaluated novel mechanisms for in-memory processing and batching that enable the stateful processing of CEP operators even under high rates of ingested events. Our evaluation demonstrates that CEPless can be easily integrated into existing CEP systems like Apache Flink while attaining similar throughput under high scale of events (up to 100K events per second) and dynamic operator update in ~238 ms.
Manisha Luthra, Sebastian Hennig, Kamran Razavi, Lin Wang 0015, Boris Koldehofe
IEEE BigData4
2020 PStream: Priority-Based Stream Scheduling for Heterogeneous Paths in Multipath-QUIC
abstract
Web latency remains the main obstacle to improving user experience with the continuous development of the web. A lot of works have been made in this course. Quick UDP Internet Connection (QUIC) embeds stream multiplexing to solve the head-of-line blocking caused by the in-order requirement of TCP. Multipath-QUIC (MPQUIC) brings further improvements by utilizing multiple paths, as is done in MultiPath TCP (MPTCP). Different from MPTCP schedulers, MPQUIC schedulers are stream-aware and thus can provide finer granularity of multipath scheduling. As streams are with different features based on their contents, the resource preferences of a stream are highly related to its feature. We find that scheduling without the recognition of the stream features can aggravate inter-stream blocking when sharing paths. We fill this gap and propose PStream - a priority-based online stream scheduling mechanism for MPQUIC, which performs path scheduling based on the stream features. We examine the effectiveness of PStream under different path heterogeneity comparing to the original and the latest scheduler of MPQUIC. Our evaluation shows that our scheduler can reduce up to 25.4% of page load time in high path heterogeneity.
Lin Wang 0015, Fa Zhang 0001, Biyu Zhou, Zhiyong Liu 0002
ICCCN2
2020 Clownfish: Edge and Cloud Symbiosis for Video Stream Analytics
abstract
Deep learning (DL) has shown promising results on complex computer vision tasks for video stream analytics recently. However, DL-based analytics typically requires intensive computation, which imposes challenges to the current computing infrastructure. In particular, cloud-only solutions struggle to maintain stable real-time performance due to the streaming over the best-effort Internet, while edge-only solutions require the DL model to be optimized (e.g., pruned or quantized) carefully to fit on resource-constrained devices, affecting the analytics quality. In this paper, we propose Clownfish, a framework for efficient video stream analytics that achieves symbiosis of the edge and the cloud. Clownfish deploys a lightweight optimized DL model at the edge for fast response and a complete DL model at the cloud for high accuracy. By exploiting the temporal correlation in video content, Clownfish sends only a subset of video frames intermittently to the cloud and enhances the analytics quality by fusing the results from the cloud model with these from the edge model. Our evaluation based on a system prototype shows that Clownfish always runs in real time and is able to achieve analytics quality comparable to that of cloud-only solutions, even under highly variable network conditions. Clownfish is generally applicable to all video stream analytics tasks that can leverage temporal correlations.
Vinod Nigade, Lin Wang 0015, Henri E. Bal
SEC2
2020 Letting off STEAM: Distributed Runtime Traffic Scheduling for Service Function Chaining
abstract
Network function virtualization has introduced a high degree of flexibility for orchestrating service functions. The provisioning of chains of service functions requires making decisions on both (1) placement of service functions and (2) scheduling of traffic through them. The placement problem (1) can be tackled during the planning phase, by exploiting coarse-grained traffic information, and has been studied extensively. However, runtime traffic scheduling (2) for optimizing system utilization and service quality, as required for future edge cloud and mobile carrier scenarios, has not been addressed so far.We fill this gap by presenting a queuing-based system model to characterize the runtime traffic scheduling problem for service function chaining. We propose a throughput-optimal scheduling policy, called integer allocation maximum pressure policy (IA-MPP). To ensure practicality in large distributed settings, we propose multi-site cooperative IA-MPP (STEAM), fulfilling runtime requirements while achieving near-optimal performance. We examine our policies in various settings representing real-world scenarios. STEAM closely matches IA-MPP in terms of throughput, and significantly outperforms (possible adaptations of) existing static or coarse-grained dynamic solutions, requiring 30%-60% less server capacity for similar service quality. Our STEAM prototype shows feasibility running on a standard server.
Marcel Blöcher, Ramin Khalili, Lin Wang 0015, Patrick Eugster
INFOCOM3
2020 Eco-friendly Caching and Forwarding in Named Data Networking
abstract
Green networking, by making the network more energy efficient and helping reduce environmental impact, is receiving more and more attraction for sustainable ICT. In this paper, we propose a new green approach for Named Data Networking (NDN) where content requests and caching perform towards green content delivery. We design a forwarding and caching strategy, where we first define the greenness of nodes, a quantitative metric for measuring the environmental footprint of the network, based on which we identify corresponding green paths and encourage traffic to aggregate on green paths powered by more eco-friendly renewable energy. We validate our approach with a variety of simulations using real network topology and renewable energy datasets from the US, and the results show that applying the proposed green NDN achieves significant ecofriendly gains.
Seng-Kyoun Jo, Lin Wang 0015, Jussi Kangasharju, Max Mühlhäuser
LANMAN2
2020 Joint Relaying and Spatial Sharing Multicast Scheduling for mmWave Networks
abstract
Millimeter-wave (mmWave) communication plays a vital role in disseminating large volumes of data in beyond-5G networks efficiently. Unfortunately, the directionality of mmWave communication significantly complicates efficient data dissemination, particularly in multicasting, which is gaining more and more importance in emerging applications (e.g., V2X, public safety, massive IoT). While multicasting for systems operating at lower frequencies (i.e., sub-6GHz) has been extensively studied, they are sub-optimal for mmWave systems as mmWave has significantly different propagation characteristics, i.e., using the directional transmission to compensate for the high path loss and thus promoting spectrum sharing. In this paper, we propose novel multicast scheduling algorithms by jointly exploiting relaying and spatial sharing gains while aiming to minimize the multicast completion time. We first characterize the problem with a comprehensive model and formulate it with an integer linear program (ILP). We further design a practical and scalable semi-distributed algorithm named mmDiMu, based on gradually maximizing the transmission throughput over time. Finally, we carry out validation through extensive simulations in different scales, and the results show that mmDiMu significantly outperforms conventional algorithms with around 95% reduction on multicast completion time.
Allyson Sim, Mahdi Mousavi, Lin Wang 0015, Anja Klein 0002, Matthias Hollick
WoWMoM3
2019 Aves: A Decision Engine for Energy-efficient Stream Analytics across Low-power Devices
abstract
Today's low-power devices, such as smartphones and wearables, form a very heterogeneous ecosystem. Applications in such a system typically follow a reactive pattern based on stream analytics, i.e., sensing, processing, and actuating. Despite the simplicity of this pattern, deciding where to place the processing tasks of an application to achieve energy efficiency is non-trivial in a heterogeneous system since application components are distributed across multiple devices. In this paper, we present Aves - a decision-making engine based on a holistic energy-prediction model, with which the processing tasks of applications can be placed automatically in an energy-efficient manner without programmer/user intervention. We validate the effectiveness of the model and reveal several counter-intuitive placement decisions. Our decision engine's improvements are typically 10-30%, with up to a factor 14 in the most extreme cases. We also show that Aves gives an accurate decision in comparison with real energy measurements for two sensor-based applications.
Roshan Bharath Das, Marc X. Makkes, Alexandru Uta, Lin Wang 0015, Henri E. Bal
IEEE BigData4
2019 A Programming Framework for Heterogeneous Stream Analytics
abstract
Sensor-based applications using Big Data are of increasing importance in various fields. A typical example of such use cases is building health-care applications [1], [2]. A typical scenario is where a patient's heart rate is monitored by a smartwatch. A smartphone can then analyze the gathered data and identify patterns in the patient's heart rate. However, if the data analysis is too complex to be performed on a smartphone, the computation could be offloaded to a nearby cloudlet or a remote cloud. A decision usually follows the analysis, and actuation is performed accordingly (e.g., a message is sent to either the patient or the doctor). Developing such an application is intrinsically complex, as the programmer needs to reconcile different APIs specific to different platforms.
Roshan Bharath Das, Marc X. Makkes, Alexandru Uta, Lin Wang 0015, Henri E. Bal
IEEE BigData4
2019 A Microservice Store for Efficient Edge Offloading
abstract
Current edge computing frameworks require tight coupling between mobile clients and surrogates, i.e., the offloaded code has been preconfigured with its required execution environment. In many cases, this includes prior transfers of code blocks or execution environments from mobile devices to the offloading infrastructure. This approach incurs additional latency and is detrimental for the energy consumption of the mobile devices. In this paper, we propose the concept of a microservice store. Using the microservice abstraction common in software development and following the serverless paradigm, we envision a repository through which said services are made accessible to developers and can be re-used across applications. We implement a proof-of-concept edge computing system based on a microservice repository and demonstrate its benefits with real-world applications on mobile devices. Our results show that we were able to reduce latencies by up to 14x and save up to 94% of battery life.
Julien Gedeon, Jens Heuschkel, Lin Wang 0015, Max Mühlhäuser
GLOBECOM4
2019 Incentivizing Microservices for Online Resource Sharing in Edge Clouds
abstract
The microservice architecture provides high agility, making it a suitable choice for implementing edge cloud services. Provisioning microservices at the network edge requires the dynamic allocation of resources. However, due to the resource limitation in the edge cloud environment, there is no guarantee that enough resources are always available upon a microservice's requests. In this paper, we design an online auction-based mechanism to incentivize microservices to spare their occupied resources so that the edge cloud platform can reclaim them and reallocate them to other microservices that need resources. We firstly design a single-stage auction that determines the winning bids to satisfy the resource demands in polynomial time, while calculating the payments. Then, we design an online framework to tie a series of such single-stage auctions into a multi-stage online mechanism without requiring the knowledge of future bids and demands. Via rigorous analysis, we exhibit that our mechanism design achieves truthful bidding and individual rationality, with a constant competitive ratio regarding the social cost of the system in the long run. Finally, we verify the practical performance of our mechanism through extensive simulations.
Amit Samanta 0001, Lei Jiao 0002, Max Mühlhäuser, Lin Wang 0015
ICDCS4
2019 FStream: Flexible Stream Scheduling and Prioritizing in Multipath-QUIC
abstract
While the web keeps evolving, web latency remains a major obstacle to improving user experience. In the past, many efforts have been made in this course. SPDY achieves reduced latency through multiplexing and prioritization by manipulating HTTP. Quick UDP Internet Connection (QUIC) generalizes the idea and embeds multiplexing in the transport layer by introducing application-oriented streams. Multipath-QUIC brings further improvements by utilizing multiple paths as is done in MultiPath TCP (MPTCP). However, failing to account for stream priorities in the transport layer can result in suboptimal performance for time-critical streams. We fill this gap and propose FStream - a flexible stream scheduling mechanism for Multipath-QUIC, which provides stream prioritization down to the transport layer. We implement FStream in Multipath-QUIC and demonstrate its effectiveness in reducing the completion time of time-critical streams ( 3x) through extensive experiments under different path dissimilarity conditions.
Lin Wang 0015, Fa Zhang 0001, Zhiyong Liu 0002
ICPADS2
2019 An Online Market Mechanism for Edge Emergency Demand Response via Cloudlet Control
abstract
The computing frontier is moving from centralized mega datacenters towards distributed cloudlets at the network edge. We argue that cloudlets are well-suited for participation in Emergency Demand Response (EDR) programs due to their enormous energy consumption and flexible workload distribution, while existing EDR mechanisms for clouds and colocation datacenters are not suitable for cloudlets. We propose a novel online market mechanism, EdgeEDR, to incentivize cloudlets to participate in EDR, featuring multiple cloudlet-specific designs. At a high level, we observe that cloudlet operators can dynamically switch on/off entire cloudlets to compensate for the energy reduction required by the power grid. We formulate a long-term social cost minimization problem and decompose it into a series of one-round procurement auctions. In each auction instance, we propose to let the cloudlet tenants bid with cost functions of their service quality degradation tolerance, and let the cloudlet operator choose the service quality, allocate the workload, and shut down the cloudlets. Via rigorous analysis, we exhibit that our bidding policy is individually rational and truthful; our workload distribution algorithm has near-optimal performance in each auction; and our overall online algorithm achieves a provable competitive ratio. We further confirm the performance of our mechanism through extensive trace-driven simulations.
Lei Jiao 0002, Lin Wang 0015, Fangming Liu
INFOCOM3
2019 PABO: Mitigating congestion via packet bounce in data center networks
Lin Wang 0015, Fa Zhang 0001, Kai Zheng 0003, Max Mühlhäuser, Zhiyong Liu 0002
Comput. Commun.2
2019 MOERA: Mobility-Agnostic Online Resource Allocation for Edge Computing
abstract
To better support emerging interactive mobile applications such as those VR-/AR-based, cloud computing is quickly evolving into a new computing paradigm called edge computing. Edge computing has the promise of bringing cloud resources to the network edge to augment the capability of mobile devices in close proximity to the user. One big challenge in edge computing is the efficient allocation and adaptation of edge resources in the presence of high dynamics imposed by user mobility. This paper provides a formal study of this problem. By characterizing a variety of static and dynamic performance measures with a comprehensive cost model, we formulate the online edge resource allocation problem with a mixed nonlinear optimization problem. We propose MOERA, a mobility-agnostic online algorithm based on the “regularization” technique, which can be used to decompose the problem into separate subproblems with regularized objective functions and solve them using convex programming. Through rigorous analysis we are able to prove that MOERA can guarantee a parameterized competitive ratio, without requiring any a priori knowledge on input. We carry out extensive experiments with various real-world data and show that MOERA can achieve an empirical competitive ratio of less than 1.2, reduces the total cost by $4 \times$4× compared to static approaches, and outperforms the online greedy one-shot solution by 70 percent. Moreover, we verify that even being future-agnostic, MOERA can achieve comparable performance to approaches with perfect partial future knowledge. We also discuss practical issues with respect to the implementation of our algorithm in real edge computing systems.
Lin Wang 0015, Lei Jiao 0002, Jun Li 0001, Julien Gedeon, Max Mühlhäuser
IEEE Trans. Mob. Comput.1
2018 On Scalable In-Network Operator Placement for Edge Computing
abstract
The drawbacks encountered in today's cloud computing infrastructures have led to a paradigm shift towards in-network processing, where resources in the core and at the edge of the network are leveraged to perform computations. This can lead to decreased costs and better quality of service for users, e.g., when latency-critical applications are executed close to data sources and users. Deploying applications or parts thereof on these infrastructures requires to place operators (i.e., functional components of applications) on available resources in the network. Solving large instances of this problem in an optimal way is known to be computationally hard and, thus, practically unfeasible. While heuristic approaches exist, they mostly aim at placing functionalities on homogeneous nodes or make unrealistic assumptions for edge computing environments. To address this issue, this paper studies the placement problem in the context of a 3-tier architecture consisting of cloud, fog and edge devices. We provide a comprehensive model and propose a heuristic approach to the problem, in which we introduce constraints on the placement decision to limit the possible solution space, leading to a decrease in the solving time for the problem. These constraints exploit the characteristics of our 3-tier network architecture. To demonstrate the feasibility of the approach, we present a general framework that supports different types of heuristics. We validate the approach by implementing example heuristics for each type. We show that our approach can scale to large instances, i.e., it can significantly reduce the resolution time to find a placement solution while introducing only a small optimality gap.
Julien Gedeon, Michael Stein 0001, Lin Wang 0015, Max Mühlhäuser
ICCCN3
2018 Cost-Effective and Eco-Friendly Green Routing Using Renewable Energy
abstract
While communication technologies are evolving rapidly, there is still the nontrivial matter of a communication systems being green. Although some energy-aware solutions have been proposed for the telecommunications sector, they are not designed with the ultimate goal of being environment-friendly. In this paper, we investigate the problem of achieving energy efficiency in IP networks by taking into account not only the energy consumption of the network but also the impact of various energy sources, e.g., renewable energies. We propose a new green networking approach in which we classify network nodes into clusters and select one header node in each cluster according to the generation cost and the carbon emission per unit of energy. We develop a routing scheme using IP routing only on header nodes and conducting packet forwarding using a carefully designed identifier on other nodes to achieve a greener communication system. We validate our solution with a variety of simulations using real-world renewable energy statistics, and the results show that our approach is superior to other existing solutions, particularly in terms of energy and cost efficiency.
Seng-Kyoun Jo, Lin Wang 0015, Jussi Kangasharju, Max Mühlhäuser
ICCCN2
2018 Service Entity Placement for Social Virtual Reality Applications in Edge Computing
abstract
While social Virtual Reality (VR) applications such as Facebook Spaces are becoming popular, they are not compatible with classic mobile-or cloud-based solutions due to their processing of tremendous data and exchange of delay-sensitive metadata. Edge computing may fulfill these demands better, but it is still an open problem to deploy social VR applications in an edge infrastructure while supporting economic operations of the edge clouds and satisfactory quality-of-service for the users. This paper presents the first formal study of this problem. We model and formulate a combinatorial optimization problem that captures all intertwined goals. We propose ITEM, an iterative algorithm with fast and big “moves” where in each iteration, we construct a graph to encode all the costs and convert the cost optimization into a graph cut problem. By obtaining the minimum s-t cut via existing max-flow algorithms, we can simultaneously determine the placement of multiple service entities, and thus, the original problem can be addressed by solving a series of graph cuts. Our evaluations with large-scale, real-world data traces demonstrate that ITEM converges fast and outperforms baseline approaches by more than 2 × in one-shot placement and around 1.3 × in dynamic, online scenarios where users move arbitrarily in the system.
Lin Wang 0015, Lei Jiao 0002, Ting He 0001, Jun Li 0001, Max Mühlhäuser
INFOCOM1
2018 VirtualStack: Flexible Cross-layer Optimization via Network Protocol Virtualization
abstract
The world is driven by the Internet and there is no doubt about its importance in our daily life. However, the Internet has rarely been upgraded since its advent, although the ISO OSI model has already provided the required flexibility. With innovations being blocked, the Internet is suffering from a high-degree of ossification (e.g., the slow progress of IPv6 update), leading to suboptimal efficiency for emerging applications as well as enlarged maintenance cost.In this paper, we present VirtualStack, which aims at bringing back the interchangeability of network layers. VirtualStack is based on the idea of protocol virtualization, where the most suitable protocol stack can be dynamically composed and applied on the fly according to the characteristics of both the application and the physical link. Through a comprehensive study, we show that many existing but not widely deployed protocols outperform the omnipresent TCP under various link technologies and network conditions. This provides the necessary insight for dynamic composition of the network protocol stack. We further evaluate VirtualStack under a typical Internet setting with multiple hops under different conditions. The experimental results confirm the benefits as well as the potential of VirtualStack.
Jens Heuschkel, Lin Wang 0015, Erik Fleckstein, Michael Ofenloch, Marcel Blöcher, Jon Crowcroft, Max Mühlhäuser
LCN2
2018 Multiple Granularity Online Control of Cloudlet Networks for Edge Computing
abstract
Operating distributed cloudlets at optimal cost is nontrivial when facing not only the dynamic and unpredictable resource prices and user requests, but also the low efficiency of today's immature cloudlet infrastructures. We propose to control cloudlet networks at multiple granularities: fine-grained control of servers inside cloudlets and coarse-grained control of cloudlets themselves. We model this problem as a mixed-integer nonlinear program with the switching cost over time. To solve this problem online, we firstly linearize, "regularize", and decouple it into a series of one-shot subproblems that we solve at each corresponding time slot, and afterwards we design an iterative, dependent rounding framework using our proposed randomized pairwise rounding algorithm to convert the fractional control decisions into the integral ones at each time slot. Via rigorous theoretical analysis, we exhibit our approach's performance guarantee in terms of the competitive ratio and the multiplicative integrality gap towards the offline optimal integral decisions. Extensive evaluations with real-world data confirm the empirical superiority of our approach over the single granularity server control and the state-of-the-art algorithms.
Lei Jiao 0002, Lingjun Pu, Lin Wang 0015, Xiaojun Lin 0001, Jun Li 0001
SECON3
2018 Online Resource Allocation, Content Placement and Request Routing for Cost-Efficient Edge Caching in Cloud Radio Access Networks
abstract
In this paper, we advocate edge caching in cloud radio access networks (C-RAN) to facilitate the ever-increasing mobile multimedia services. In our framework, central offices will cooperatively allocate cloud resources to cache popular contents and satisfy user requests for those contents, so as to minimize the system costs in terms of storage, VM reconfiguration, content access latency, and content migration. However, this joint resource allocation, content placement and request routing, is nontrivial, since it needs to be continuously adjusted to accommodate system dynamics, such as user movement and content slashdot effect, while taking into account the time-correlated adjustment costs for VM reconfiguration and content migration. To this end, we build a comprehensive model to capture the key components of edge caching in C-RAN and formulate a joint optimization problem, aiming at minimizing the system costs over time and meanwhile satisfying the time-varying user requests and respecting various practical constraints (e.g., storage and bandwidth). Then, we propose a novel online approximation algorithm by resorting to the regularization, rounding, and decomposition technique, which can be proved to have a parameterized competitive ratio with a polynomial running time. Extensive trace-driven simulations corroborate the efficiency, flexibility, and lightweight of our proposed online algorithm; for instance, it achieves an empirical competitive ratio around 2 - 4 and gains over 30% improvement compared with many state-of-the-art algorithms in various system settings.
Lingjun Pu, Lei Jiao 0002, Xu Chen 0004, Lin Wang 0015, Qinyi Xie, Jingdong Xu
IEEE J. Sel. Areas Commun.4
2017 Distributed Graph-based Topology Adaptation using Motif Signatures
abstract
A motif is a small graph pattern, and a motif signature counts the occurrences of selected motifs in a network. The motif signature of a real-world network is an important characteristic because it is closely related to a variety of semantic and functional aspects. In recent years, motif analysis has been successfully applied for adapting topologies of communication networks: The motif signatures of very good networks (e.g., in terms of load balancing) are determined a priori to derive a target motif signature. Then, a given network is adapted in iterative steps, subject to side constraints and in a distributed way, such that its motif signature approximates the target motif signature. In this paper, we formalize this adaptation problem and show that it is 풩풫-hard. We present LoMbA, a generic approach for motif-based graph adaptation: All types of networks, all selections of motifs, and all types of consistency-maintaining constraints can be incorporated. To evaluate LoMbA, we conduct a simulation study based on several scenarios of topology adaptation from the domain of communication networks. We consider topology control in wireless ad-hoc networks, balancing of video streaming trees, and load balancing of peer-to-peer overlays. In each considered application scenario, the simulation results are remarkably good, although the implementation was not tuned toward these scenarios.
Michael Stein 0001, Karsten Weihe, Augustin Wilberg, Roland Speith, Julian M. Klomp, Mathias Schnee, Lin Wang 0015, Max Mühlhäuser
ALENEX7
2017 Joint Optimization of Server and Network Resource Utilization in Cloud Data Centers
abstract
Virtual machine placement is a key component of cloud resource management, which may affect network bandwidth allocation. In this paper, we revisit the virtual machine placement problem in cloud data centers and aim to maximize the overall resource utilization in multiple dimensions, while ensuring that the resource constraints on both the server such as CPU capacity and the network such as bandwidth are not violated. We model the bandwidth-guaranteed virtual machine placement problem and prove its NP-hardness, and design offline and online algorithms to solve the problem. We first consider the offline version and develop approximation algorithms with bounded performance ratios for both the homogeneous and the heterogeneous cases. Then, for the online version, we propose simple and efficient heuristics based on the insights from the offline algorithm design. Comprehensive experimental results verify that the overall resource utilization can be significantly improved by applying our proposals.
Biyu Zhou, Jie Wu 0001, Lin Wang 0015, Fa Zhang 0001, Zhiyong Liu 0002
GLOBECOM3
2017 PABO: Congestion mitigation via packet bounce
abstract
Today's data center applications can generate a diverse mix of short and long flows. However, switches used in a typical data center network are usually shallow buffered in order to reduce queueing delay and deployment cost. As a result, the buildup of the queues by long flows can block short flows, leading to frequent packet losses and retransmissions, which translates to crucial performance degradation. While multiple end-to-end TCP-based solutions have been proposed, none of them have tackled the real challenge: reliable transmission in the network. In this paper, we fill this gap by presenting PABO — a novel link-layer design that can mitigate congestion by temporarily bouncing packets to upstream switches. PABO's design fulfills the following demands: i) providing per-flow based flow control on the link layer, ii) handling transient congestion without the intervention of end devices, and iii) gradually back propagating the congestion signal to the source when the network is not capable to handle the congestion. We complete a proof-of-concept implementation, and experiments under different severities of congestion show that PABO outperforms the standard unreliable link-layer protocol by guaranteeing zero packet loss while introducing only a reasonable stretch on packet delay.
Lin Wang 0015, Fa Zhang 0001, Kai Zheng 0003, Zhiyong Liu 0002
ICC2
2017 Online Resource Allocation for Arbitrary User Mobility in Distributed Edge Clouds
abstract
As clouds move to the network edge to facilitate mobile applications, edge cloud providers are facing new challenges on resource allocation. As users may move and resource prices may vary arbitrarily, %and service delays are heterogeneous, resources in edge clouds must be allocated and adapted continuously in order to accommodate such dynamics. In this paper, we first formulate this problem with a comprehensive model that captures the key challenges, then introduce a gap-preserving transformation of the problem, and propose a novel online algorithm that optimally solves a series of subproblems with a carefully designed logarithmic objective, finally producing feasible solutions for edge cloud resource allocation over time. We further prove via rigorous analysis that our online algorithm can provide a parameterized competitive ratio, without requiring any a priori knowledge on either the resource price or the user mobility. Through extensive experiments with both real-world and synthetic data, we further confirm the effectiveness of the proposed algorithm. We show that the proposed algorithm achieves near-optimal results with an empirical competitive ratio of about 1.1, reduces the total cost by up to 4x compared to static approaches, and outperforms the online greedy one-shot optimizations by up to 70%.
Lin Wang 0015, Lei Jiao 0002, Jun Li 0001, Max Mühlhäuser
ICDCS1
2017 Online Flow Scheduling with Deadline for Energy Conservation in Data Center Networks
abstract
We study the problem of flow scheduling in data center networks. Using speed scaling, our aim is to find an online scheduling algorithm that minimizes the total energy consumption of the network by determining both the transmission order and rates of the arriving flows while providing a strict flow deadline guarantee. Observing the superlinear property of link power consumption, the key challenge is in constantly determining the minimum transmission rate for “delay-tolerable” flows without any priori knowledge. To leverage the flow arrival pattern, we propose a probability-based flow prediction model to capture the uncertainty of the network flows. Based on the prediction model, we propose a tunable online flow scheduling algorithm to solve the online flow scheduling problem effectively. By introducing a scaling factor on bandwidth allocation, this algorithm allows us to conduct arbitrary trade-offs between the conservative and aggressive behaviors in terms of energy conser- vation. The effectiveness of the proposed algorithm is validated through rigorous theoretical analysis and further confirmed by extensive numerical simulations.
Biyu Zhou, Jie Wu 0001, Lin Wang 0015, Fa Zhang 0001, Zhiyong Liu 0002
ICPADS3
2017 Green routing using renewable energy for IP networks
abstract
Green technology for not only reducing energy consumption but also environmental pollution has become a critical factor in ICT industries. However, for the telecommunications sector in particular, most network elements are not usually optimized for power efficiency. In this work, we propose a green routing method in an IP network for the reduction of unnecessary energy consumption. In addition, it can encourage the use of power generated by renewable energy sources instead of using traditional fossil energy. As a green networking approach, we first classify the network nodes into either header or member nodes according to the quantity of the available renewable energies. The member nodes then put the routing related module at layer 3 to sleep based on the assumption that this layer in the OSI model can operate independently. All of the network nodes are then partitioned into clusters consisting of one header node and multiple member nodes. Then, only the header node in a cluster conducts IP routing and its member nodes conduct packet switching using a specially designed identifier, referred to as a tag. To investigate the impact of the proposed scheme, we conducted a number of simulations using real-world renewable energy statistics and results show that our approach outperforms the existing solutions in terms of energy efficiency to a large extent.
Seng-Kyoun Jo, Lin Wang 0015, Max Mühlhäuser, Jussi Kangasharju
LANMAN2
2017 Identifying the Performance Impairment of HTTP
abstract
Online web services have been constantly developed and offered through the Internet. These services rely on webpages that are hosted on remote servers and are transmitted to clients via HTTP upon requests. While HTTP has been updated to support more and more sophisticated services with increasing amount of multimedia content, the real-world adoption of those updates is still very slow. In this paper, we investigate the performance impairment in current HTTP implementations and suggest solutions to improve the situation. Specifically, we identify two insights, i.e., avoiding unnecessary DNS lookups by replacing domain names with prefetched IP addresses for linked resources and reducing TCP connections by attaching linked resources to HTML files as binary stream, that could be employed to largely improve the performance of web services without modifying HTTP itself. Through extensive measurements, we validate our insights and the results confirm our findings.
Jens Heuschkel, Jens Forstmann, Lin Wang 0015, Max Mühlhäuser
LCN3
2017 Joint Optimization of Operational Cost and Performance Interference in Cloud Data Centers
abstract
Virtual machine (VM) scheduling is an important technique for the efficient operation of the computing resources in a data center. Previous work has mainly focused on consolidating VMs to improve resource utilization and to optimize energy consumption. However, the interference between collocated VMs is usually ignored, which can result in much worse performance degradation of the applications running on the VMs due to the contention of the shared resources. Based on this observation, we aim at designing efficient VM assignment and scheduling strategies in which we consider optimizing both the operational cost of the data center and the performance degradation of the running applications. We then propose a general model that captures the tradeoff between the two contradictory objectives. We present offline and online solutions for this problem by exploiting the spatial and temporal information of performance interference of VM collocation, where VM scheduling is performed by jointly considering the combinations and the life-cycle overlap of the VMs. Evaluation results show that the proposed methods can generate efficient schedules for VMs, achieving low operational cost while significantly reducing the performance degradation of applications in cloud data centers.
Xibo Jin, Fa Zhang 0001, Lin Wang 0015, Songlin Hu 0001, Biyu Zhou, Zhiyong Liu 0002
IEEE Trans. Cloud Comput.3
2017 LazyCtrl: A Scalable Hybrid Network Control Plane Design for Cloud Data Centers
abstract
The advent of software defined networking enables flexible, reliable and feature-rich control planes for data center networks. However, the tight coupling of centralized control and complete visibility leads to a wide range of issues among which scalability has risen to prominence due to the excessive workload on the central controller. By analyzing the traffic patterns from a couple of production data centers, we observe that data center traffic is usually highly skewed and thus edge switches can be clustered into a set of communication-intensive groups according to traffic locality. Motivated by this observation, we present LazyCtrl, a novel hybrid control plane design for data center networks where network control is carried out by distributed control mechanisms inside independent groups of switches while complemented with a global controller. LazyCtrl aims at bringing laziness to the global controller by dynamically devolving most of the control tasks to independent switch groups to process frequent intra-group events near the datapath while handling rare inter-group or other specified events by the controller. We implement LazyCtrl and build a prototype based on Open vSwitch and Floodlight. Trace-driven experiments on our prototype show that an effective switch grouping is easy to maintain in multi-tenant clouds and the central controller can be significantly shielded by staying “lazy”, with its workload reduced by up to 82 percent.
Kai Zheng 0003, Lin Wang 0015, Baohua Yang, Yi Sun 0004, Steve Uhlig
IEEE Trans. Parallel Distributed Syst.2
2016 Reconciling task assignment and scheduling in mobile edge clouds
abstract
The prosperous growth of the Internet-of-Things industry attracts numerous interests in employing edge clouds (a.k.a. cloudlets) to enhance the performance of mobile services and applications. Most existing research has been focused on offloading computational tasks from mobile devices to a single cloudlet or a central location, yet overlooked the issue of jointly coordinating the offloaded tasks in a system of multiple cloudlets. In this paper, we fill this gap by investigating the assignment and the scheduling of mobile computational tasks over multiple cloudlets, while optimizing the overall cost efficiency by leveraging the heterogeneity of cloudlets. We model both data transfer and computation in terms of monetary and time costs, with task deadlines guaranteed. We formulate the problem as a mixed integer program and prove its NP-hardness. By introducing admission control for the cloudlet provider to shape the system workload, we transform our problem into maximizing the task admission rate over the two coupled phases: data transfer and computation. We propose an efficient two-phase scheduling algorithm, and demonstrate that, compared with the conventional approach of always selecting the closest cloudlet, our approach achieves significantly higher admission rate with up to 20% reduction in the average cost of all offloaded tasks.
Lin Wang 0015, Lei Jiao 0002, Dzmitry Kliazovich, Pascal Bouvry
ICNP1
2016 Power-efficient assignment of virtual machines to physical machines
Jordi Arjona Aroca, Antonio Fernández 0001, Miguel A. Mosteiro, Christopher Thraves, Lin Wang 0015
Future Gener. Comput. Syst.5
2016 HDEER: A Distributed Routing Scheme for Energy-Efficient Networking
abstract
The proliferation of new online Internet services has substantially increased the energy consumption in wired networks, which has become a critical issue for Internet service providers. In this paper, we target the network-wide energy-saving problem by leveraging speed scaling as the energy-saving strategy. We propose a distributed routing scheme-HDEER-to improve network energy efficiency in a distributed manner without significantly compromising traffic delay. HDEER is a two-stage routing scheme where a simple distributed multipath finding algorithm is firstly performed to guarantee loop-free routing, and then a distributed routing algorithm is executed for energy-efficient routing in each node among the multiple loop-free paths. We conduct extensive experiments on the NS3 simulator and simulations with real network topologies in different scales under different traffic scenarios. Experiment results show that HDEER can reduce network energy consumption with a fair tradeoff between network energy consumption and traffic delay.
Biyu Zhou, Fa Zhang 0001, Lin Wang 0015, Chenying Hou, Antonio Fernández 0001, Athanasios V. Vasilakos, Youshi Wang, Jie Wu 0001, Zhiyong Liu 0002
IEEE J. Sel. Areas Commun.3
2015 Lazy Ctrl: Scalable Network Control for Cloud Data Centers
abstract
The advent of software defined networking enables flexible, reliable and feature-rich control planes for data center networks. However, the tight coupling of centralized control and complete visibility leads to a wide range of issues among which scalability has risen to prominence. We observe that data center traffic is usually highly skewed and thus edge switches can be grouped according to traffic locality. As a result, the workload of the central controller could be highly reduced if we carry out distributed control inside those groups. Based on the above observation, we present LazyCtrl, a novel hybrid control plane design for data center networks. LazyCtrl aims at bringing laziness to the central controller by dynamically devolving most of the control tasks to independent switch groups to process frequent intra-group events using distributed control mechanisms, while handling rare inter-group or other specified events by the controller. We implement LazyCtrl and build a prototype based on Open vSwich and Floodlight. Trace-driven experiments on our prototype show that an effective switch grouping is easy to maintain in multi-tenant clouds and the central controller can be significantly shielded by staying lazy, with its workload reduced by up to 82%.
Lin Wang 0015, Kai Zheng 0003, Baohua Yang, Yi Sun 0004, Steve Uhlig
ICDCS1
2015 PROP: Using PCIe-Based RDMA to Accelerate Rack-Scale Communications in Data Centers
abstract
In order to reduce the demands on bandwidth of core layer network, data center operators usually assign tasks of the same job to servers that are located in the same rack, leading to the fact that 80% of the traffic originated from servers retains in the same rack. As a result, providing sufficient network capacity inside racks becomes critical to the Quality-of-Service of current data center applications. In this paper, we propose PROP, a novel hybrid network architecture which leverages PCIe-based RDMA to reinforce rack-scale connectivity in data centers. In our design, intra-rack bulk data transfers will be accelerated by a dedicated high-bandwidth PCIe-compliant network while complemented with the existing Ethernet network. In addition, we develop a proprietary PCIe-based RDMA hardware which can allow the servers in the same rack to exchange data in main memory without involving the operating system and the processors. We also implement a software stack to enable existing socket-based applications to transparently utilize the proposed dedicated network system. As the preliminary stage, this paper focuses on exploiting the unique design point and implements an FPGA-based prototype to validate the technical feasibility of the proposed architecture.
Dawei Zang, Zheng Cao 0003, Xiaoli Liu 0002, Lin Wang 0015, Zhan Wang 0003, Ninghui Sun
ICPADS4
2015 Multi-resource energy-efficient routing in cloud data centers with network-as-a-service
abstract
With the rapid development of software defined networking and network function virtualization, researchers have proposed a new cloud networking model called Network-as-a-Service (NaaS) which enables both in-network packet processing and application-specific network control. In this paper, we revisit the problem of achieving network energy efficiency in data centers and identify some new optimization challenges under the NaaS model. Particularly, we extend the energy-efficient routing optimization from single-resource to multi-resource settings. We characterize the problem through a detailed model and provide a formal problem definition. Due to the high complexity of direct solutions, we propose a greedy routing scheme to approximate the optimum, where flows are selected progressively to exhaust residual capacities of active nodes, and routing paths are assigned based on the distributions of both node residual capacities and flow demands. By leveraging the structural regularity of data center networks, we also provide a fast topology-aware heuristic method based on hierarchically solving a series of vector bin packing instances. Extensive simulations show that the proposed routing scheme can achieve significant gain on energy savings and the topology-aware heuristic can produce comparably good results while reducing the computation time to a large extent.
Lin Wang 0015, Antonio Fernández 0001, Fa Zhang 0001, Jie Wu 0001, Zhiyong Liu 0002
ISCC1
2014 Energy-Efficient Flow Scheduling and Routing with Hard Deadlines in Data Center Networks
abstract
The power consumption of enormous network devices in data centers has emerged as a big concern to data center operators. Despite many traffic-engineering-based solutions, very little attention has been paid on performance-guaranteed energy saving schemes. In this paper, we propose a novel energy-saving model for data center networks by scheduling and routing "deadline-constrained flows" where the transmission of every flow has to be accomplished before a rigorous deadline, being the most critical requirement in production data center networks. Based on speed scaling and power-down energy saving strategies for network devices, we aim to explore the most energy efficient way of scheduling and routing flows on the network, as well as determining the transmission speed for every flow. We consider two general versions of the problem. For the version of only flow scheduling where routes of flows are pre-given, we show that it can be solved polynomially and we develop an optimal combinatorial algorithm for it. For the version of joint flow scheduling and routing, we prove that it is strongly NP-hard and cannot have a Fully Polynomial-Time Approximation Scheme (FPTAS) unless P=NP. Based on a relaxation and randomized rounding technique, we provide an efficient approximation algorithm which can guarantee a provable performance ratio with respect to a polynomial of the total number of flows.
Lin Wang 0015, Fa Zhang 0001, Kai Zheng 0003, Athanasios V. Vasilakos, Shaolei Ren, Zhiyong Liu 0002
ICDCS1
2014 DEER: A distributed routing scheme for achieving network energy efficiency
abstract
The rapid growth of Internet services has brought emergent concerns over network energy efficiency. This study aims to improve network energy efficiency using power-down technique. We propose DEER, a fully distributed routing scheme. The main concept of DEER is to dynamically allocate the traffic demands in the nodes so that some links connected to the nodes can be put into sleep mode, thus reducing the energy consumption. The special features of DEER include that it does not need global traffic matrix of the network and that it uses only the local information of link loads, making DEER be able to be implemented in a distributed manner without centralized control. We develop algorithms in DEER to dynamically change the link state (into active or sleep mode) according to link utilization and to balance the loads of links by adjusting the link weights. With the traffic load varying over time, the link state transformation is triggered when any of the pre-defined thresholds is violated. Extensive simulations with the network topology and real traffic traces from the GÉANT network confirm that by involving DEER, up to 50% of the links can be put into sleep while the frequency of chaining the state of a link stays fairly low.
Biyu Zhou, Lin Wang 0015, Fa Zhang 0001, Xibo Jin, Zhiyong Liu 0002
LANMAN2
2014 GreenDCN: A General Framework for Achieving Energy Efficiency in Data Center Networks
abstract
The popularization of cloud computing has raised concerns over the energy consumption that takes place in data centers. In addition to the energy consumed by servers, the energy consumed by large numbers of network devices emerges as a significant problem. Existing work on energy-efficient data center networking primarily focuses on traffic engineering, which is usually adapted from traditional networks. We propose a new framework to embrace the new opportunities brought by combining some special features of data centers with traffic engineering. Based on this framework, we characterize the problem of achieving energy efficiency with a time-aware model, and we prove its NP-hardness with a solution that has two steps. First, we solve the problem of assigning virtual machines (VM) to servers to reduce the amount of traffic and to generate favorable conditions for traffic engineering. The solution reached for this problem is based on three essential principles that we propose. Second, we reduce the number of active switches and balance traffic flows, depending on the relation between power consumption and routing, to achieve energy conservation. Experimental results confirm that, by using this framework, we can achieve up to 50 percent energy savings. We also provide a comprehensive discussion on the scalability and practicability of the framework.
Lin Wang 0015, Fa Zhang 0001, Jordi Arjona Aroca, Athanasios V. Vasilakos, Kai Zheng 0003, Chenying Hou, Dan Li 0001, Zhiyong Liu 0002
IEEE J. Sel. Areas Commun.1
2013 Improving the Network Energy Efficiency in MapReduce Systems
abstract
Apart from servers, the energy consumed by enormous amount of network devices in data centers also emerges as a big problem. Existing work on energy- efficient data center networking primarily focuses on traffic engineering to consolidate flows and shut down unused devices, not considering another important factor, virtual machine assignment, which has been shown to have a big influence on traffic engineering. Moreover, the lack of information about upper layer applications leads to misunderstand the traffic patterns of the network. This may result in poor effectiveness in the traffic-based optimization in practice. In this paper, we aim to achieve better network energy efficiency in MapReduce systems by combining virtual machine assignment and traffic engineering. By exploiting the characteristics of MapReduce applications, we provide a unified model to describe this problem. Due to its NP-hardness, a general framework is proposed to solve it, where virtual machines are first clustered and then different virtual machine assignments are generated greedily and a local search procedure is used to improve them. The local search procedure depends on the results of an energy-efficient routing provided by GEERA. GEERA is an approximate algorithm designed to select routing paths for flows. Experimental results confirm the efficiency of GEERA, as well as the overall framework. By using this framework, up to $20\%$ more energy savings can be achieved compared with sole traffic engineering solutions.
Lin Wang 0015, Fa Zhang 0001, Zhiyong Liu 0002
ICCCN1
2013 Incorporating Rate Adaptation Into Green Networking for Future Data Centers
abstract
Despite some proposals for energy-efficient topologies, most of the studies for saving energy in data center networks are focused on traffic engineering, i.e., consolidating flows and switching off unnecessary network devices. The major weakness of this approach is network oscillation brought by the frequent change of network topology when traffic fluctuates very fast. In this paper, we propose to incorporate rate adaptation into green data center networks. With rate adaptive network devices, we aim at approaching network-wide energy proportionality by routing optimization. We formalize the problem with an integer program and propose an efficient approximation algorithm - TSRR, solving the problem quickly while guaranteeing a constant performance ratio. Extensive range of simulations confirm that more than 40% of the energy can be saved while introducing very slight stretch on network delay.
Lin Wang 0015, Fa Zhang 0001, Chenying Hou, Jordi Arjona Aroca, Zhiyong Liu 0002
NCA1
2012 Energy-Efficient Network Routing with Discrete Cost Functions
Lin Wang 0015, Antonio Fernández 0001, Fa Zhang 0001, Chenying Hou, Zhiyong Liu 0002
TAMC1