Ming Zhang 0005

dblp:73/1844-5 · DBLP profile ↗
← Back
56ranked-venue papers
5as first author
6since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 46 · 2 first-author · 6 since 2021Systems, architecture and hardware · 7 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-authorSecurity and privacy · 1
YearPublicationVenuePosition
2023 Buffer-Based High-Coverage and Low-Overhead Request Event Monitoring in the Cloud
abstract
Request latency directly affects the performance of modern cloud applications. Due to various causes in hosts and networks, requests can suffer from request latency anomalies (RLAs), which may violate the Service-Level Agreement. However, existing performance monitoring tools have incomplete coverage and inconsistent semantics for monitoring requests and cannot accurately diagnose RLAs. This paper presentsBufScope, a high-coverage and low-overhead request event monitoring system, which monitorsbuffersto capture most RLA-related abnormal events with consistent request-level semantics in the end-to-end datapath of request. First,BufScopemodels the datapath of request as a buffer chain and defines events based on three properties of buffers, so as toend-to-end monitorthe root causes of RLA. Then, to achieveconsistent semanticsfor captured events,BufScopedesigns a request-level semantics injection mechanism to make events captured in networks have the victim requests’ ID. Finally,BufScopeoffloads the semantics operations and event collection in software to SmartNICs forlow CPU overhead. We have implementedBufScopeon commodity SmartNICs and programmable switches. Evaluation results show thatBufScopecan diagnose 98% RLAs with < 0.08% network bandwidth overhead and 0.6% application throughput decline.
Kaihui Gao, Chen Sun 0005, Shuai Wang 0028, Dan Li 0001, Yu Zhou 0008, Hongqiang Harry Liu, Lingjun Zhu, Ming Zhang 0005, Lu Lu 0016
IEEE/ACM Trans. Netw.8
2022 Buffer-based End-to-end Request Event Monitoring in the Cloud
Kaihui Gao, Chen Sun 0005, Shuai Wang 0028, Dan Li 0001, Yu Zhou 0008, Hongqiang Harry Liu, Lingjun Zhu, Ming Zhang 0005
NSDI8
2022 Predictable vFabric on informative data plane
abstract
In multi-tenant data centers, each tenant desires reassuring predictability from the virtual network fabric - bandwidth guarantee, work conservation, and bounded tail latency. Achieving these goals simultaneously relies on rapid and precise traffic admission. However, the slow convergence (tens of milliseconds) of prior works can hardly satisfy the increasingly rigorous performance demand under dynamic traffic patterns. Further, state-of-the-art load balance schemes are all guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works.
Shuai Wang 0028, Kaihui Gao, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Ming Zhang 0005
SIGCOMM16
2022 GSO-simulcast: global stream orchestration in simulcast video conferencing systems
abstract
We present GSO-Simulcast, a new architecture designed for large-scale multi-party video-conferencing systems. GSO-Simulcast is currently deployed at full-scale in Alibaba's Dingtalk video conferencing that serves more than 500 million users. It marks a fundamental shift from today's Simulcast, where a media server locally decides how to switch and forward video streams based on a fragmented network view. Instead, GSO-Simulcast globally orchestrates the publishing, subscribing, as well as the resolution and bitrate of video streams for each participant using a centralized controller that is aware of all network constraints in a meeting. The controller automatically modifies stream configurations to meet the participants' real-time network changes and updates. In doing so, GSO-Simulcast achieves multiple goals: (1) reducing video and network mismatch, (2) less path congestion, and (3) automated stream policy management. With the deployment of GSO-Simulcast, we observed more than a 35% reduction in the average video stall, 50% reduction in the average voice stall, and 6% improvement in the average video framerate. We describe the principle, design, deployment, and lessons learned.
Xianshang Lin, Junshao Zhang, Yao Cui, Dennis Cai, Hongqiang Harry Liu, Ming Zhang 0005
SIGCOMM10
2021 Aquila: a practically usable verification system for production-scale programmable data planes
abstract
This paper presents Aquila, the first practically usable verification system for Alibaba's production-scale programmable data planes. Aquila addresses four challenges in building a practically usable verification: (1) specification complexity; (2) verification scalability; (3) bug localization; and (4) verifier self validation. Specifically, first, Aquila proposes a high-level language that facilitates easy expression of specifications, reducing lines of specification codes by tenfold compared to the state-of-the-art. Second, Aquila constructs a sequential encoding algorithm to circumvent the exponential growth of states associated with the upscaling of data plane programs to production level. Third, Aquila adopts an automatic and accurate bug localization approach that can narrow down suspects based on reported violations and pinpoint the culprit by simulating a fix for each suspect. Fourth and finally, Aquila can perform self validation based on refinement proof, which involves the construction of an alternative representation and subsequent equivalence checking. To this date, Aquila has been used in the verification of our production-scale programmable edge networks for over half a year, and it has successfully prevented many potential failures resulting from data plane bugs.
Bingchuan Tian, Mengqi Liu 0001, Ennan Zhai, Yu Zhou 0008, Mengjing Ma, Xionglie Wei, Hongqiang Harry Liu, Ming Zhang 0005, Chen Tian 0001, Minlan Yu
SIGCOMM14
2021 XLINK: QoE-driven multi-path QUIC transport in large-scale video services
abstract
We report XLINK, a multi-path QUIC video transport solution with experiments in Taobao short videos. XLINK is designed to meet two operational challenges at the same time: (1) Optimized user-perceived quality of experience (QoE) in terms of robustness, smoothness, responsiveness, and mobility and (2) Minimized cost overhead for service providers (typically CDNs). The core of XLINK is to take the opportunity of QUIC as a user-space protocol and directly capture user-perceived video QoE intent to control multi-path scheduling and management. We overcome major hurdles such as multi-path head-of-line blocking, network heterogeneity, and rapid link variations and balance cost and performance.
Zhilong Zheng, Yanmei Liu, Furong Yang, Zhenyu Li 0001, Yuanbo Zhang, Jiuhai Zhang, Qing An, Hai Hong, Hongqiang Harry Liu, Ming Zhang 0005
SIGCOMM14
2020 Flow Event Telemetry on Programmable Data Plane
abstract
Network performance anomalies (NPAs), e.g. long-tailed latency, bandwidth decline, etc., are increasingly crucial to cloud providers as applications are getting more sensitive to performance. The fundamental difficulty to quickly mitigate NPAs lies in the limitations of state-of-the-art network monitoring solutions --- coarse-grained counters, active probing, or packet telemetry either cannot provide enough insights on flows or incur too much overhead. This paper presents NetSeer, a flow event telemetry (FET) monitor which aims to discover and record all performance-critical data plane events, e.g. packet drops, congestion, path change, and packet pause. NetSeer is efficiently realized on the programmable data plane. It has a high coverage on flow events including inter-switch packet drop/corruption which is critical but also challenging to retrieve the original flow information, with novel intra- and inter-switch event detection algorithms running on data plane; NetSeer also achieves high scalability and accuracy with innovative designs of event aggregation, information compression, and message batching that mainly run on data plane, using switch CPU as complement. NetSeer has been implemented on commodity programmable switches and NICs. With real case studies and extensive experiments, we show NetSeer can reduce NPA mitigation time by 61%-99% with only 0.01% overhead of monitoring traffic.
Yu Zhou 0008, Chen Sun 0005, Hongqiang Harry Liu, Rui Miao 0001, Bo Li 0061, Zhilong Zheng, Lingjun Zhu, Yongqing Xi, Dennis Cai, Ming Zhang 0005, Mingwei Xu 0001
SIGCOMM13
2020 Lyra: A Cross-Platform Language and Compiler for Data Plane Programming on Heterogeneous ASICs
abstract
Programmable data plane has been moving towards deployments in data centers as mainstream vendors of switching ASICs enable programmability in their newly launched products, such as Broadcom's Trident-4, Intel/Barefoot's Tofino, and Cisco's Silicon One. However, current data plane programs are written in low-level, chip-specific languages (e.g., P4 and NPL) and thus tightly coupled to the chip-specific architecture. As a result, it is arduous and error-prone to develop, maintain, and composite data plane programs in production networks. This paper presents Lyra, the first cross-platform, high-level language & compiler system that aids the programmers in programming data planes efficiently. Lyra offers a one-big-pipeline abstraction that allows programmers to use simple statements to express their intent, without laboriously taking care of the details in hardware; Lyra also proposes a set of synthesis and optimization techniques to automatically compile this "big-pipeline" program into multiple pieces of runnable chip-specific code that can be launched directly on the individual programmable switches of the target network. We built and evaluated Lyra. Lyra not only generates runnable real-world programs (in both P4 and NPL), but also uses up to 87.5% fewer hardware resources and up to 78% fewer lines of code than human-written programs.
Ennan Zhai, Hongqiang Harry Liu, Rui Miao 0001, Yu Zhou 0008, Bingchuan Tian, Chen Sun 0005, Dennis Cai, Ming Zhang 0005, Minlan Yu
SIGCOMM9
2020 Accuracy, Scalability, Coverage: A Practical Configuration Verifier on a Global WAN
abstract
This paper presents Hoyan-- the first reported large scale deployment of configuration verification in a global-scale wide area network (WAN). Hoyan has been running in production for more than two years and is currently used for all critical configuration auditing and updates on the WAN. We highlight our innovative designs and real-life experience to make Hoyan accurate and scalable in practice. For accuracy under the inconsistencies of devices' vendor-specific behaviors (VSBs), Hoyan continuously discovers the flaws in device behavior models, thus aiding the operators in fixing the models. For scalability to verify our global WAN, Hoyan introduces a "global-simulation & local formal-modeling" strategy to model uncertainties in small scales and perform aggressive pruning of possibilities during the protocol simulations. Hoyan achieves near-100% verification accuracy after it detected and fixed O(10) VSBs on our WAN. Hoyan has prevented many potential service failures resulting from misconfiguration and reduced the failure rate of updates of our WAN by more than half in 2019.
Fangdan Ye, Ennan Zhai, Hongqiang Harry Liu, Bingchuan Tian, Qiaobo Ye, Chunsheng Wang, Tianchen Guo, Duncheng She, Biao Cheng, Ming Zhang 0005, Rodrigo Fonseca
SIGCOMM15
2020 NFC+: Breaking NFC Networking Limits through Resonance Engineering
abstract
Current UHF RFID systems suffer from two long-standing problems: 1) miss-reading non-line-of-sight or misoriented tags and 2) cross-reading undesired, distant tags due to multi-path reflections. This paper proposes a novel system, NFC+, to overcome the fundamental challenges. NFC+ is a magnetic field reader, which can inventory standard NFC tagged objects with a reasonably long range and arbitrary orientation. NFC+ achieves this by leveraging physical and algorithmic techniques based on magnetic resonance engineering. We build a prototype of NFC+ and conduct extensive evaluations in a logistic network. Comparing to UHF RFID, we find that NFC+ can reduce the miss-reading rate from 23% to 0.03%, and cross-reading rate from 42% to 0, for randomly oriented objects. NFC+ demonstrates high robustness for RFID unfriendly media (e.g., water bottles and metal cans). It can reliably read commercial NFC tags at a distance of up to 3 meters which, for the first time, enables NFC to be directly applied to practical logistics network applications.
Renjie Zhao 0001, Purui Wang, Hongqiang Harry Liu, Xianshang Lin, Xinyu Zhang 0003, Chenren Xu, Ming Zhang 0005
SIGCOMM9
2019 HPCC: high precision congestion control
abstract
Congestion control (CC) is the key to achieving ultra-low latency, high bandwidth and network stability in high-speed networks. From years of experience operating large-scale and high-speed RDMA networks, we find the existing high-speed CC schemes have inherent limitations for reaching these goals. In this paper, we present HPCC (High Precision Congestion Control), a new high-speed CC mechanism which achieves the three goals simultaneously. HPCC leverages in-network telemetry (INT) to obtain precise link load information and controls traffic precisely. By addressing challenges such as delayed INT information during congestion and overreac-tion to INT information, HPCC can quickly converge to utilize free bandwidth while avoiding congestion, and can maintain near-zero in-network queues for ultra-low latency. HPCC is also fair and easy to deploy in hardware. We implement HPCC with commodity programmable NICs and switches. In our evaluation, compared to DCQCN and TIMELY, HPCC shortens flow completion times by up to 95%, causing little congestion even under large-scale incasts.
Rui Miao 0001, Hongqiang Harry Liu, Lingbo Tang, Zheng Cao 0003, Ming Zhang 0005, Frank Kelly, Mohammad Alizadeh, Minlan Yu
SIGCOMM8
2019 Safely and automatically updating in-network ACL configurations with intent language
abstract
In-network Access Control List (ACL) is an important technique in ensuring network-wide connectivity and security. As cloud-scale WANs today constantly evolve in size and complexity, in-network ACL rules are becoming increasingly more complex. This presents a great challenge to the updating process of ACL configurations: network operators are frequently required to update "tangled" ACL rules across thousands of devices to meet diverse business requirements, and even a single ACL misconfiguration may lead to network disruptions. Such increasing challenges call for an automated system to improve the efficiency and correctness of ACL updates. This paper presents Jinjing, a system that aids Alibaba's network operators in automatically and correctly updating ACL configurations in Alibaba's global WAN. Jinjing allows the operators to express in a declarative language, named LAI, their update intent (e.g., ACL migration and traffic control). Then, Jinjing automatically synthesizes ACL update plans that satisfy their intent. At the heart of Jinjing, we develop a set of novel verification and synthesis techniques to rigorously guarantee the correctness of update plans. In Alibaba, our operators have used Jinjing to efficiently update their ACLs and have thus prevented significant service downtime.
Bingchuan Tian, Xinyi Zhang 0003, Ennan Zhai, Hongqiang Harry Liu, Qiaobo Ye, Chunsheng Wang, Zhiming Ji, Yihong Sang, Ming Zhang 0005, Chen Tian 0001, Haitao Zheng 0001, Ben Y. Zhao
SIGCOMM10
2017 CherryPick: Adaptively Unearthing the Best Cloud Configurations for Big Data Analytics
Omid Alipourfard, Hongqiang Harry Liu, Jianshu Chen, Shivaram Venkataraman, Minlan Yu, Ming Zhang 0005
NSDI6
2017 Guaranteeing Deadlines for Inter-Data Center Transfers
abstract
Inter-data center wide area networks (inter-DC WANs) carry a significant amount of data transfers that require to be completed within certain time periods, or deadlines. However, very little work has been done to guarantee such deadlines. The crux is that the current inter-DC WAN lacks an interface for users to specify their transfer deadlines and a mechanism for provider to ensure the completion while maintaining high WAN utilization. In this paper, we address the problem by introducing a deadline-based network abstraction (DNA) for inter-DC WANs. DNA allows users to explicitly specify the amount of data to be delivered and the deadline by which it has to be completed. The malleability of DNA provides flexibility in resource allocation. Based on this, we develop a system calledAmoebathat implements DNA. Our simulations and test bed experiments show thatAmoeba, by harnessing DNA’s malleability, accommodates 15% more user requests with deadlines, while achieving 60% higher WAN utilization than prior solutions.
Hong Zhang 0025, Kai Chen 0005, Wei Bai 0001, Dongsu Han, Chen Tian 0001, Hao Wang 0022, Haibing Guan, Ming Zhang 0005
IEEE/ACM Trans. Netw.8
2016 Yoda: a highly available layer-7 load balancer
abstract
Layer-7 load balancing is a foundational building block of online services. The lack of offerings from major public cloud providers have left online services to build their own load balancers (LB), or use third-party LB design such as HAProxy. The key problem with such proxy-based design is each proxy instance is a single point of failure, as upon its failure, the TCP flow state for the connections with the client and server is lost which breaks the user flows. This significantly affects user experience and online services revenue.
Rohan Gandhi, Y. Charlie Hu, Ming Zhang 0005
EuroSys3
2016 Efficiently Delivering Online Services over Integrated Infrastructure
Hongqiang Harry Liu, Raajay Viswanathan, Matt Calder, Aditya Akella, Ratul Mahajan, Jitendra Padhye, Ming Zhang 0005
NSDI7
2015 Guaranteeing deadlines for inter-datacenter transfers
abstract
Inter-datacenter wide area networks (inter-DC WAN) carry a significant amount of data transfers that require to be completed within certain time periods, or deadlines. However, very little work has been done to guarantee such deadlines. The crux is that the current inter-DC WAN lacks an interface for users to specify their transfer deadlines and a mechanism for provider to ensure the completion while maintaining high WAN utilization.
Hong Zhang 0025, Kai Chen 0005, Wei Bai 0001, Dongsu Han, Chen Tian 0001, Hao Wang 0022, Haibing Guan, Ming Zhang 0005
EuroSys8
2015 Congestion Control for Large-Scale RDMA Deployments
abstract
Modern datacenter applications demand high throughput (40Gbps) and ultra-low latency (< 10 μs per hop) from the network, with low CPU overhead. Standard TCP/IP stacks cannot meet these requirements, but Remote Direct Memory Access (RDMA) can. On IP-routed datacenter networks, RDMA is deployed using RoCEv2 protocol, which relies on Priority-based Flow Control (PFC) to enable a drop-free network. However, PFC can lead to poor application performance due to problems like head-of-line blocking and unfairness. To alleviates these problems, we introduce DCQCN, an end-to-end congestion control scheme for RoCEv2. To optimize DCQCN performance, we build a fluid model, and provide guidelines for tuning switch buffer thresholds, and other protocol parameters. Using a 3-tier Clos network testbed, we show that DCQCN dramatically improves throughput and fairness of RoCEv2 RDMA traffic. DCQCN is implemented in Mellanox NICs, and is being deployed in Microsoft's datacenters.
Yibo Zhu 0001, Haggai Eran, Daniel Firestone, Chuanxiong Guo, Marina Lipshteyn, Yehonatan Liron, Jitendra Padhye, Shachar Raindel, Mohamad Haj Yahia, Ming Zhang 0005
SIGCOMM10
2015 Packet-Level Telemetry in Large Datacenter Networks
abstract
Debugging faults in complex networks often requires capturing and analyzing traffic at the packet level. In this task, datacenter networks (DCNs) present unique challenges with their scale, traffic volume, and diversity of faults. To troubleshoot faults in a timely manner, DCN administrators must a) identify affected packets inside large volume of traffic; b) track them across multiple network components; c) analyze traffic traces for fault patterns; and d) test or confirm potential causes. To our knowledge, no tool today can achieve both the specificity and scale required for this task.
Yibo Zhu 0001, Nanxi Kang, Jiaxin Cao, Albert G. Greenberg, Guohan Lu, Ratul Mahajan, David A. Maltz, Ming Zhang 0005, Ben Y. Zhao, Haitao Zheng 0001
SIGCOMM9
2015 Rubik: Unlocking the Power of Locality and End-point Flexibility in Cloud Scale Load Balancing
Rohan Gandhi, Y. Charlie Hu, Cheng-Kok Koh, Hongqiang Harry Liu, Ming Zhang 0005
USENIX ATC5
2014 Duet: cloud scale load balancing with hardware and software
abstract
Load balancing is a foundational function of datacenter infrastructures and is critical to the performance of online services hosted in datacenters. As the demand for cloud services grows, expensive and hard-to-scale dedicated hardware load balancers are being replaced with software load balancers that scale using a distributed data plane that runs on commodity servers. Software load balancers offer low cost, high availability and high flexibility, but suffer high latency and low capacity per load balancer, making them less than ideal for applications that demand either high throughput, or low latency or both. In this paper, we present Duet, which offers all the benefits of software load balancer, along with low latency and high availability -- at next to no cost. We do this by exploiting a hitherto overlooked resource in the data center networks -- the switches themselves. We show how to embed the load balancing functionality into existing hardware switches, thereby achieving organic scalability at no extra cost. For flexibility and high availability, Duet seamlessly integrates the switch-based load balancer with a small deployment of software load balancer. We enumerate and solve several architectural and algorithmic challenges involved in building such a hybrid load balancer. We evaluate Duet using a prototype implementation, as well as extensive simulations driven by traces from our production data centers. Our evaluation shows that Duet provides 10x more capacity than a software load balancer, at a fraction of a cost, while reducing latency by a factor of 10 or more, and is able to quickly adapt to network dynamics including failures.
Rohan Gandhi, Hongqiang Harry Liu, Y. Charlie Hu, Guohan Lu, Jitendra Padhye, Ming Zhang 0005
SIGCOMM7
2014 Dynamic scheduling of network updates
abstract
We present Dionysus, a system for fast, consistent network updates in software-defined networks. Dionysus encodes as a graph the consistency-related dependencies among updates at individual switches, and it then dynamically schedules these updates based on runtime differences in the update speeds of different switches. This dynamic scheduling is the key to its speed; prior update methods are slow because they pre-determine a schedule, which does not adapt to runtime conditions. Testbed experiments and data-driven simulations show that Dionysus improves the median update speed by 53--88% in both wide area and data center networks compared to prior methods.
Xin Jin 0008, Hongqiang Harry Liu, Rohan Gandhi, Srikanth Kandula, Ratul Mahajan, Ming Zhang 0005, Jennifer Rexford, Roger Wattenhofer
SIGCOMM6
2014 Traffic engineering with forward fault correction
abstract
Faults such as link failures and high switch configuration delays can cause heavy congestion and packet loss. Because it takes time to detect and react to faults, these conditions can last long---even tens of seconds. We propose forward fault correction (FFC), a proactive approach to handling faults. FFC spreads network traffic such that freedom from congestion is guaranteed under arbitrary combinations of up to k faults. We show how FFC can be practically realized by compactly encoding the constraints that arise from this large number of possible faults and solving them efficiently using sorting networks. Experiments with data from real networks show that, with negligible loss in overall network throughput, FFC can reduce data loss by a factor of 7--130 in well-provisioned networks, and reduce the loss of high-priority traffic to almost zero in well-utilized networks.
Hongqiang Harry Liu, Srikanth Kandula, Ratul Mahajan, Ming Zhang 0005, David Gelernter
SIGCOMM4
2014 A network-state management service
abstract
We present Statesman, a network-state management service that allows multiple network management applications to operate independently, while maintaining network-wide safety and performance invariants. Network state captures various aspects of the network such as which links are alive and how switches are forwarding traffic. Statesman uses three views of the network state. In observed state, it maintains an up-to-date view of the actual network state. Applications read this state and propose state changes based on their individual goals. Using a model of dependencies among state variables, Statesman merges these proposed states into a target state that is guaranteed to maintain the safety and performance invariants. It then updates the network to the target state. Statesman has been deployed in ten Microsoft Azure datacenters for several months, and three distinct applications have been built on it. We use the experience from this deployment to demonstrate how Statesman enables each application to meet its goals, while maintaining network-wide invariants.
Ratul Mahajan, Jennifer Rexford, Ming Zhang 0005, Ahsan Arefin
SIGCOMM5
2013 Achieving high utilization with software-driven WAN
abstract
We present SWAN, a system that boosts the utilization of inter-datacenter networks by centrally controlling when and how much traffic each service sends and frequently re-configuring the network's data plane to match current traffic demand. But done simplistically, these re-configurations can also cause severe, transient congestion because different switches may apply updates at different times. We develop a novel technique that leverages a small amount of scratch capacity on links to apply updates in a provably congestion-free manner, without making any assumptions about the order and timing of updates at individual switches. Further, to scale to large networks in the face of limited forwarding table capacity, SWAN greedily selects a small set of entries that can best satisfy current demand. It updates this set without disrupting traffic by leveraging a small amount of scratch capacity in forwarding tables. Experiments using a testbed prototype and data-driven simulations of two production networks show that SWAN carries 60% more traffic than the current practice.
Chi-Yao Hong, Srikanth Kandula, Ratul Mahajan, Ming Zhang 0005, Vijay Gill, Mohan Nanduri, Roger Wattenhofer
SIGCOMM4
2013 zUpdate: updating data center networks with zero loss
abstract
Datacenter networks (DCNs) are constantly evolving due to various updates such as switch upgrades and VM migrations. Each update must be carefully planned and executed in order to avoid disrupting many of the mission-critical, interactive applications hosted in DCNs. The key challenge arises from the inherent difficulty in synchronizing the changes to many devices, which may result in unforeseen transient link load spikes or even congestions. We present one primitive, zUpdate, to perform congestion-free network updates under asynchronous switch and traffic matrix changes. We formulate the update problem using a network model and apply our model to a variety of representative update scenarios in DCNs. We develop novel techniques to handle several practical challenges in realizing zUpdate as well as implement the zUpdate prototype on OpenFlow switches and deploy it on a testbed that resembles real DCN topology. Our results, from both real-world experiments and large-scale trace-driven simulations, show that zUpdate can effectively perform congestion-free updates in production DCNs.
Hongqiang Harry Liu, Ming Zhang 0005, Roger Wattenhofer, David A. Maltz
SIGCOMM3
2013 Rake: Semantics Assisted Network-Based Tracing Framework
abstract
The ability to trace request execution paths is critical for diagnosing performance faults in large-scale distributed systems. Previous black-box and white-box approaches are either inaccurate or invasive. We present a novel semantics-assisted gray-box tracing approach, called Rake, which can accurately trace individual request by observing network traffic. Rake infers the causality between messages by identifying polymorphic IDs in messages according to application semantics. To make Rake universally applicable, we design a Rake language so that users can easily describe necessary semantics of their applications while reusing the core Rake component. We evaluate Rake using a few popular distributed applications, including web search, distributed computing cluster, content provider network, and online chatting. Our results demonstrate Rake is much more accurate than the black-box approaches while requiring no modification to OS/applications. In the CoralCDN (a content distributed network) experiments, Rake links messages with much higher accuracy than WAP5, a state-of-the-art black-box approach. In the Hadoop (a distributed computing cluster platform) experiments, Rake helps reveal several previously unknown issues that may lead to performance degradation, including a RPC (Remote Procedure Call) abusing problem.
Yao Zhao 0003, Yinzhi Cao, Yan Chen 0004, Ming Zhang 0005, Anup Goyal
IEEE Trans. Netw. Serv. Manag.4
2012 Where is the energy spent inside my app?: fine grained energy accounting on smartphones with Eprof
abstract
Where is the energy spent inside my app? Despite the immense popularity of smartphones and the fact that energy is the most crucial aspect in smartphone programming, the answer to the above question remains elusive. This paper first presents eprof, the first fine-grained energy profiler for smartphone apps. Compared to profiling the runtime of applications running on conventional computers, profiling energy consumption of applications running on smartphones faces a unique challenge, asynchronous power behavior, where the effect on a component's power state due to a program entity lasts beyond the end of that program entity. We present the design, implementation and evaluation of eprof on two mobile OSes, Android and Windows Mobile.
Abhinav Pathak, Y. Charlie Hu, Ming Zhang 0005
EuroSys3
2012 You Can Run, but You Can't Hide: Exposing Network Location for Targeted DoS Attacks in Cellular Networks
Zhiyun Qian, Zhaoguang Wang, Z. Morley Mao, Ming Zhang 0005, Yi-Min Wang
NDSS5
2012 NetPilot: automating datacenter network failure mitigation
abstract
Driven by the soaring demands for always-on and fast-response online services, modern datacenter networks have recently undergone tremendous growth. These networks often rely on commodity hardware to reach immense scale while keeping capital expenses under check. The downside is that commodity devices are prone to failures, raising a formidable challenge for network operators to promptly handle these failures with minimal disruptions to the hosted services.
Daniel Turner, Chao-Chih Chen, David A. Maltz, Xiaowei Yang 0001, Ming Zhang 0005
SIGCOMM7
2011 MicroTE: fine grained traffic engineering for data centers
abstract
The effects of data center traffic characteristics on data center traffic engineering is not well understood. In particular, it is unclear how existing traffic engineering techniques perform under various traffic patterns, namely how do the computed routes differ from the optimal routes. Our study reveals that existing traffic engineering techniques perform 15% to 20% worse than the optimal solution. We find that these techniques suffer mainly due to their inability to utilize global knowledge about flow characteristics and make coordinated decision for scheduling flows.
Theophilus Benson, Ashok Anand, Aditya Akella, Ming Zhang 0005
CoNEXT4
2011 Fine-grained power modeling for smartphones using system call tracing
abstract
Accurate, fine-grained online energy estimation and accounting of mobile devices such as smartphones is of critical importance to understanding and debugging the energy consumption of mobile applications. We observe that state-of-the-art, utilization-based power modeling correlates the (actual) utilization of a hardware component with its power state, and hence is insufficient in capturing several power behavior not directly related to the component utilization in modern smartphones. Such behavior arise due to various low level power optimizations programmed in the device drivers. We propose a new, system-call-based power modeling approach which gracefully encompasses both utilization-based and non-utilization-based power behavior. We present the detailed design of such a power modeling scheme and its implementation on Android and Windows Mobile. Our experimental results using a diverse set of applications confirm that the new model significantly improves the fine-grained as well as whole-application energy consumption accuracy. We further demonstrate fine-grained energy accounting enabled by such a fined-grained power model, via amanually implemented eprof, the energy counterpart of the classic gprof tool, for profiling application energy drain.
Abhinav Pathak, Y. Charlie Hu, Ming Zhang 0005, Paramvir Bahl, Yi-Min Wang
EuroSys3
2011 Bootstrapping energy debugging on smartphones: a first look at energy bugs in mobile devices
abstract
This paper argues that a new class of bugs faced by millions of smartphones, energy bugs or ebugs, have become increasingly prominent that already they have led to significant user frustrations. We take a first look at this emerging important technical challenge faced by the smartphones, ebugs, broadly defined as an error in the system (application, OS, hardware, firmware, external conditions or combination) that causes an unexpected amount of high energy consumption by the system as a whole. We first present a taxonomy of the kinds of ebugs based on mining over 39K posts (1.2M before filtering) from 4 online mobile user forum and mobile OS bug repositories. The taxonomy shows the highly diverse nature of smartphone ebugs. We then propose a roadmap towards developing a systematic diagnosing framework for debugging ebugs on smartphones.
Abhinav Pathak, Y. Charlie Hu, Ming Zhang 0005
HotNets3
2011 Latency inflation with MPLS-based traffic engineering
abstract
While MPLS has been extensively deployed in recent years, little is known about its behavior in practice. We examine the performance of MPLS in Microsoft's online service network (MSN), a well-provisioned multi-continent production network connecting tens of data centers. Using detailed traces collected over a 2-month period, we find that many paths experience significantly inflated latencies. We correlate occurrences of latency inflation with routers, links, and DC-pairs. This analysis sheds light on the causes of latency inflation and suggests several avenues for alleviating the problem.
Abhinav Pathak, Ming Zhang 0005, Y. Charlie Hu, Ratul Mahajan, David A. Maltz
Internet Measurement Conference2
2011 Rake: Semantics assisted network-based tracing framework
abstract
The ability to trace request execution paths is critical for diagnosing performance faults in large-scale distributed systems. Previous black-box and white-box approaches are either inaccurate or invasive. We present a novel semantics-assisted gray-box tracing approach, called Rake, which can accurately trace individual request by observing network traffic. Rake infers the causality between messages by identifying polymorphic IDs in messages according to application semantics. To make Rake universally applicable, we design a Rake language so that users can easily describe necessary semantics of their applications while reusing the core Rake component. We evaluate Rake using a few popular distributed applications, including web search, distributed computing cluster, content provider network, and online chatting. Our results demonstrate Rake is much more accurate than the black-box approaches while requiring no modification to OS/applications. In the CoralCDN (a content distributed network) experiments, Rake links messages with much higher accuracy than WAP5, a state-of-the-art black-box approach. In the Hadoop (a distributed computing cluster platform) experiments, Rake helps reveal several previously unknown issues that may lead to performance degradation, including a RPC (Remote Procedure Call) abusing problem.
Yao Zhao 0003, Yinzhi Cao, Yan Chen 0004, Ming Zhang 0005, Anup Goyal
IWQoS4
2011 Switchboard: a matchmaking system for multiplayer mobile games
abstract
Supporting interactive, multiplayer games on mobile phones over cellular networks is a difficult problem. It is particularly relevant now with the explosion of mostly single-player or turn-based games on mobile phones. The challenges stem from the highly variable performance of cellular networks and the need for scalability (not burdening the cellular infrastructure, nor any server resources that a game developer deploys). We have built a service for matchmaking in mobile games -- assigning players to games such that game settings are satisfied as well as latency requirements for an enjoyable game. This requires solving two problems. First, the service needs to know the cellular network latency between game players. Second, the service needs to quickly group players into viable game sessions. In this paper, we present the design of our service, results from our experiments on predicting cellular latency, and results from efficiently grouping players into games.
Justin Manweiler, Sharad Agarwal, Ming Zhang 0005, Romit Roy Choudhury, Paramvir Bahl
MobiSys3
2011 CloudProphet: towards application performance prediction in cloud
abstract
Choosing the best-performing cloud for one's application is a critical problem for potential cloud customers. We propose CloudProphet, a trace-and-replay tool to predict a legacy application's performance if migrated to a cloud infrastructure. CloudProphet traces the workload of the application when running locally, and replays the same workload in the cloud for prediction. We discuss two key technical challenges in designing CloudProphet, and some preliminary results using a prototype implementation.
Ang Li 0002, Xuanran Zong, Srikanth Kandula, Xiaowei Yang 0001, Ming Zhang 0005
SIGCOMM5
2011 An untold story of middleboxes in cellular networks
abstract
The use of cellular data networks is increasingly popular as network coverage becomes more ubiquitous and many diverse user-contributed mobile applications become available. The growing cellular traffic demand means that cellular network carriers are facing greater challenges to provide users with good network performance and energy efficiency, while protecting networks from potential attacks. To better utilize their limited network resources while securing the network and protecting client devices the carriers have already deployed various network policies that influence traffic behavior. Today, these policies are mostly opaque, though they directly impact application designs and may even introduce network vulnerabilities.
Zhaoguang Wang, Zhiyun Qian, Z. Morley Mao, Ming Zhang 0005
SIGCOMM5
2010 CloudCmp: comparing public cloud providers
abstract
While many public cloud providers offer pay-as-you-go computing, their varying approaches to infrastructure, virtualization, and software services lead to a problem of plenty. To help customers pick a cloud that fits their needs, we develop CloudCmp, a systematic comparator of the performance and cost of cloud providers. CloudCmp measures the elastic computing, persistent storage, and networking services offered by a cloud along metrics that directly reflect their impact on the performance of customer applications. CloudCmp strives to ensure fairness, representativeness, and compliance of these measurements while limiting measurement cost. Applying CloudCmp to four cloud providers that together account for most of the cloud customers today, we find that their offered services vary widely in performance and costs, underscoring the need for thoughtful provider selection. From case studies on three representative cloud applications, we show that CloudCmp can guide customers in selecting the best-performing provider for their applications.
Ang Li 0002, Xiaowei Yang 0001, Srikanth Kandula, Ming Zhang 0005
Internet Measurement Conference4
2010 Anatomizing application performance differences on smartphones
abstract
The use of cellular data networks is increasingly popular due to the widespread deployment of 3G technologies and the rapid adoption of smartphones, such as iPhone and GPhone. Besides email and web browsing, a variety of network applications are now available, rendering smartphones potentially useful substitutes for their desktop counterparts. Nevertheless, the performance of smartphone applications in the wild is still poorly understood due to a lack of systematic measurement methodology.
Junxian Huang 0001, Birjodh Singh Tiwana, Z. Morley Mao, Ming Zhang 0005, Paramvir Bahl
MobiSys5
2010 WebProphet: Automating Performance Prediction for Web Services
Zhichun Li, Ming Zhang 0005, Zhaosheng Zhu, Yan Chen 0004, Albert G. Greenberg, Yi-Min Wang
NSDI2
2010 Optimizing Cost and Performance in Online Service Provider Networks
Zheng Zhang 0009, Ming Zhang 0005, Albert G. Greenberg, Y. Charlie Hu, Ratul Mahajan, Blaine Christian
NSDI2
2009 Detecting traffic differentiation in backbone ISPs with NetPolice
abstract
Traffic differentiations are known to be found at the edge of the Internet in broadband ISPs and wireless carriers [13, 2]. The ability to detect traffic differentiations is essential for customers to develop effective strategies for improving their application performance. We build a system, called NetPolice, that enables detection of content- and routing-based differentiations in backbone ISPs. NetPolice is easy to deploy since it only relies on loss measurement launched from end hosts. The key challenges in building NetPolice include selecting an appropriate set of probing destinations and ensuring the robustness of detection results to measurement noise.
Ying Zhang 0022, Z. Morley Mao, Ming Zhang 0005
Internet Measurement Conference3
2008 Ascertaining the Reality of Network Neutrality Violation in Backbone ISPs
Ying Zhang 0022, Z. Morley Mao, Ming Zhang 0005
HotNets3
2008 Uncovering Performance Differences Among Backbone ISPs with Netdiff
Ratul Mahajan, Ming Zhang 0005, Lindsey Poole, Vivek S. Pai
NSDI2
2008 Effective Diagnosis of Routing Disruptions from End Systems
Ying Zhang 0022, Z. Morley Mao, Ming Zhang 0005
NSDI3
2008 Automating Network Application Dependency Discovery: Experiences, Limitations, and New Solutions
Xu Chen 0028, Ming Zhang 0005, Z. Morley Mao, Paramvir Bahl
OSDI2
2007 Towards highly reliable enterprise network services via inference of multi-level dependencies
abstract
Localizing the sources of performance problems in large enterprise networks is extremely challenging. Dependencies are numerous, complex and inherently multi-level, spanning hardware and software components across the network and the computing infrastructure. To exploit these dependencies for fast, accurate problem localization, we introduce an Inference Graph model, which is well-adapted to user-perceptible problems rooted in conditions giving rise to both partial service degradation and hard faults. Further, we introduce the Sherlock system to discover Inference Graphs in the operational enterprise, infer critical attributes, and then leverage the result to automatically detect and localize problems. To illuminate strengths and limitations of the approach, we provide results from a prototype deployment in a large enterprise network, as well as from testbed emulations and simulations. In particular, we find that taking into account multi-level structure leads to a 30% improvement in fault localization, as compared to two-level approaches.
Paramvir Bahl, Ranveer Chandra, Albert G. Greenberg, Srikanth Kandula, David A. Maltz, Ming Zhang 0005
SIGCOMM6
2006 Discovering Dependencies for Network Management
Paramvir Bahl, Paul Barham 0001, Richard Black, Ranveer Chandra, Moisés Goldszmidt, Rebecca Isaacs, Srikanth Kandula, John MacCormick, David A. Maltz, Richard Mortier, Michal Wawrzoniak, Ming Zhang 0005
HotNets13
2006 WiFiProfiler: cooperative diagnosis in wireless LANs
abstract
While 802.11-based wireless hotspots are proliferating, users often have little recourse when the network does not work or performs poorly for them. They are left trying to manually debug the problem, which can be a frustrating and disruptive process. The users' troubles are compounded by the absence of network administrators or an IT department to turn to in many 802.11 hotspot settings (e.g., cafes, airports, conferences).We present WiFiProfiler, a system in which wireless hosts cooperate to diagnose and possibly resolve network problems in an automated manner, without requiring any infrastructural support. The key observation is that even if a host's wireless link to an access point is not working, the host is often within the range of other wireless nodes and is in a position to communicate with them (a little) peer-to-peer. We leverage this ability to create a shared information plane, which enables wireless hosts to exchange a range of information about their network settings and the health of their network connectivity. By aggregating and correlating such information across multiple wireless hosts, we infer the likely cause of the problem. Our implementation on Windows XP shows that WiFiProfiler is effective in diagnosing a range of problems and imposes a low overhead on the participating hosts.
Ranveer Chandra, Venkat N. Padmanabhan, Ming Zhang 0005
MobiSys3
2006 How DNS Misnaming Distorts Internet Topology Mapping
Ming Zhang 0005, Yaoping Ruan, Vivek S. Pai, Jennifer Rexford
USENIX ATC, General Track1
2004 Network-Embedded Programmable Storage and Its Applications
Sumeet Sobti, Junwen Lai, Yilei Shao, Chi Zhang 0070, Ming Zhang 0005, Fengzhou Zheng, Arvind Krishnamurthy, Randolph Y. Wang
NETWORKING6
2004 PlanetSeer: Internet Path Failure Monitoring and Characterization in Wide-Area Services
Ming Zhang 0005, Chi Zhang 0070, Vivek S. Pai, Larry L. Peterson, Randolph Y. Wang
OSDI1
2004 A Transport Layer Approach for Improving End-to-End Performance and Robustness Using Redundant Paths
Ming Zhang 0005, Junwen Lai, Arvind Krishnamurthy, Larry L. Peterson, Randolph Y. Wang
USENIX ATC, General Track1
2003 RR-TCP: A Reordering-Robust TCP with DSACK
abstract
TCP performs poorly on paths that reorder packets significantly, where it misinterprets out-of-order delivery as packet loss. The sender responds with a fast retransmit though no actual loss has occurred. These repeated false fast retransmits keep the sender's window small, and severely degrade the throughput it attains. Requiring nearly in-order delivery needlessly restricts and complicates Internet routing systems and routers. Such beneficial systems as multi-path routing and parallel packet switches are difficult to deploy in a way that preserves ordering. Toward a more reordering-tolerant Internet architecture, we present enhancements to TCP that improve the protocol's robustness to reordered and delayed packets. We extend the sender to detect and recover from false fast retransmits using DSACK information, and to avoid false fast retransmits proactively, by adaptively varying dupthresh. Our algorithm is the first that adaptively balances increasing dupthresh, to avoid false fast retransmits, and limiting the growth of dupthresh, to avoid unnecessary timeouts. Finally, we demonstrate that TCP's RTO estimator tolerates delayed packets poorly, and present enhancements to it that ensure it is sufficiently conservative, without using timestamps or additional TCP header hits. Our simulations show that these enhancements significantly improve TCP's performance over paths that reorder or delay packets.
Ming Zhang 0005, Brad Karp, Sally Floyd, Larry L. Peterson
ICNP1
2002 Probabilistic Packet Scheduling: Achieving Proportional Share Bandwidth Allocation for TCP Flows
abstract
This paper describes and evaluates a probabilistic packet scheduling (PPS) algorithm for providing different levels of service to TCP flows. With our approach, each router defines a local currency in terms of tickets and assigns tickets to its inputs based on contractual agreements with its upstream routers. A flow is tagged with tickets to represent the relative share of bandwidth it should receive at each link. When multiple flows share the same bottleneck, the bandwidth that each flow obtains is proportional to the relative tickets assigned to that flow. Simulations show that PPS does a better job of proportionally allocating bandwidth than DiffServ and weighted CSFQ. In addition, PPS accommodates flows that cross multiple currency domains.
Ming Zhang 0005, Randolph Y. Wang, Larry L. Peterson, Arvind Krishnamurthy
INFOCOM1