VLDB 2026 Research / reviewers in the wild / expert
Ming Zhang 0005
dblp:73/1844-5
· DBLP profile ↗
56ranked-venue papers
5as first author
6since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 46 · 2 first-author · 6 since 2021Systems, architecture and hardware · 7 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-authorSecurity and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Buffer-Based High-Coverage and Low-Overhead Request Event Monitoring in the CloudabstractRequest latency directly affects the performance of modern cloud applications. Due to various causes in hosts and networks, requests can suffer from request latency anomalies (RLAs), which may violate the Service-Level Agreement. However, existing performance monitoring tools have incomplete coverage and inconsistent semantics for monitoring requests and cannot accurately diagnose RLAs. This paper presentsBufScope, a high-coverage and low-overhead request event monitoring system, which monitorsbuffersto capture most RLA-related abnormal events with consistent request-level semantics in the end-to-end datapath of request. First,BufScopemodels the datapath of request as a buffer chain and defines events based on three properties of buffers, so as toend-to-end monitorthe root causes of RLA. Then, to achieveconsistent semanticsfor captured events,BufScopedesigns a request-level semantics injection mechanism to make events captured in networks have the victim requests’ ID. Finally,BufScopeoffloads the semantics operations and event collection in software to SmartNICs forlow CPU overhead. We have implementedBufScopeon commodity SmartNICs and programmable switches. Evaluation results show thatBufScopecan diagnose 98% RLAs with < 0.08% network bandwidth overhead and 0.6% application throughput decline. Kaihui Gao, Chen Sun 0005, Shuai Wang 0028, Dan Li 0001, Yu Zhou 0008, Hongqiang Harry Liu, Lingjun Zhu, Ming Zhang 0005, Lu Lu 0016 |
IEEE/ACM Trans. Netw. | 8 |
| 2022 | Buffer-based End-to-end Request Event Monitoring in the Cloud
Kaihui Gao, Chen Sun 0005, Shuai Wang 0028, Dan Li 0001, Yu Zhou 0008, Hongqiang Harry Liu, Lingjun Zhu, Ming Zhang 0005 |
NSDI | 8 |
| 2022 | Predictable vFabric on informative data planeabstractIn multi-tenant data centers, each tenant desires reassuring predictability from the virtual network fabric - bandwidth guarantee, work conservation, and bounded tail latency. Achieving these goals simultaneously relies on rapid and precise traffic admission. However, the slow convergence (tens of milliseconds) of prior works can hardly satisfy the increasingly rigorous performance demand under dynamic traffic patterns. Further, state-of-the-art load balance schemes are all guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works. Shuai Wang 0028, Kaihui Gao, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Ming Zhang 0005 |
SIGCOMM | 16 |
| 2022 | GSO-simulcast: global stream orchestration in simulcast video conferencing systemsabstractWe present GSO-Simulcast, a new architecture designed for large-scale multi-party video-conferencing systems. GSO-Simulcast is currently deployed at full-scale in Alibaba's Dingtalk video conferencing that serves more than 500 million users. It marks a fundamental shift from today's Simulcast, where a media server locally decides how to switch and forward video streams based on a fragmented network view. Instead, GSO-Simulcast globally orchestrates the publishing, subscribing, as well as the resolution and bitrate of video streams for each participant using a centralized controller that is aware of all network constraints in a meeting. The controller automatically modifies stream configurations to meet the participants' real-time network changes and updates. In doing so, GSO-Simulcast achieves multiple goals: (1) reducing video and network mismatch, (2) less path congestion, and (3) automated stream policy management. With the deployment of GSO-Simulcast, we observed more than a 35% reduction in the average video stall, 50% reduction in the average voice stall, and 6% improvement in the average video framerate. We describe the principle, design, deployment, and lessons learned. Xianshang Lin, Junshao Zhang, Yao Cui, Dennis Cai, Hongqiang Harry Liu, Ming Zhang 0005 |
SIGCOMM | 10 |
| 2021 | Aquila: a practically usable verification system for production-scale programmable data planesabstractThis paper presents Aquila, the first practically usable verification system for Alibaba's production-scale programmable data planes. Aquila addresses four challenges in building a practically usable verification: (1) specification complexity; (2) verification scalability; (3) bug localization; and (4) verifier self validation. Specifically, first, Aquila proposes a high-level language that facilitates easy expression of specifications, reducing lines of specification codes by tenfold compared to the state-of-the-art. Second, Aquila constructs a sequential encoding algorithm to circumvent the exponential growth of states associated with the upscaling of data plane programs to production level. Third, Aquila adopts an automatic and accurate bug localization approach that can narrow down suspects based on reported violations and pinpoint the culprit by simulating a fix for each suspect. Fourth and finally, Aquila can perform self validation based on refinement proof, which involves the construction of an alternative representation and subsequent equivalence checking. To this date, Aquila has been used in the verification of our production-scale programmable edge networks for over half a year, and it has successfully prevented many potential failures resulting from data plane bugs. Bingchuan Tian, Mengqi Liu 0001, Ennan Zhai, Yu Zhou 0008, Mengjing Ma, Xionglie Wei, Hongqiang Harry Liu, Ming Zhang 0005, Chen Tian 0001, Minlan Yu |
SIGCOMM | 14 |
| 2021 | XLINK: QoE-driven multi-path QUIC transport in large-scale video servicesabstractWe report XLINK, a multi-path QUIC video transport solution with experiments in Taobao short videos. XLINK is designed to meet two operational challenges at the same time: (1) Optimized user-perceived quality of experience (QoE) in terms of robustness, smoothness, responsiveness, and mobility and (2) Minimized cost overhead for service providers (typically CDNs). The core of XLINK is to take the opportunity of QUIC as a user-space protocol and directly capture user-perceived video QoE intent to control multi-path scheduling and management. We overcome major hurdles such as multi-path head-of-line blocking, network heterogeneity, and rapid link variations and balance cost and performance. Zhilong Zheng, Yanmei Liu, Furong Yang, Zhenyu Li 0001, Yuanbo Zhang, Jiuhai Zhang, Qing An, Hai Hong, Hongqiang Harry Liu, Ming Zhang 0005 |
SIGCOMM | 14 |
| 2020 | Flow Event Telemetry on Programmable Data PlaneabstractNetwork performance anomalies (NPAs), e.g. long-tailed latency, bandwidth decline, etc., are increasingly crucial to cloud providers as applications are getting more sensitive to performance. The fundamental difficulty to quickly mitigate NPAs lies in the limitations of state-of-the-art network monitoring solutions --- coarse-grained counters, active probing, or packet telemetry either cannot provide enough insights on flows or incur too much overhead. This paper presents NetSeer, a flow event telemetry (FET) monitor which aims to discover and record all performance-critical data plane events, e.g. packet drops, congestion, path change, and packet pause. NetSeer is efficiently realized on the programmable data plane. It has a high coverage on flow events including inter-switch packet drop/corruption which is critical but also challenging to retrieve the original flow information, with novel intra- and inter-switch event detection algorithms running on data plane; NetSeer also achieves high scalability and accuracy with innovative designs of event aggregation, information compression, and message batching that mainly run on data plane, using switch CPU as complement. NetSeer has been implemented on commodity programmable switches and NICs. With real case studies and extensive experiments, we show NetSeer can reduce NPA mitigation time by 61%-99% with only 0.01% overhead of monitoring traffic. Yu Zhou 0008, Chen Sun 0005, Hongqiang Harry Liu, Rui Miao 0001, Bo Li 0061, Zhilong Zheng, Lingjun Zhu, Yongqing Xi, Dennis Cai, Ming Zhang 0005, Mingwei Xu 0001 |
SIGCOMM | 13 |
| 2020 | Lyra: A Cross-Platform Language and Compiler for Data Plane Programming on Heterogeneous ASICsabstractProgrammable data plane has been moving towards deployments in data centers as mainstream vendors of switching ASICs enable programmability in their newly launched products, such as Broadcom's Trident-4, Intel/Barefoot's Tofino, and Cisco's Silicon One. However, current data plane programs are written in low-level, chip-specific languages (e.g., P4 and NPL) and thus tightly coupled to the chip-specific architecture. As a result, it is arduous and error-prone to develop, maintain, and composite data plane programs in production networks. This paper presents Lyra, the first cross-platform, high-level language & compiler system that aids the programmers in programming data planes efficiently. Lyra offers a one-big-pipeline abstraction that allows programmers to use simple statements to express their intent, without laboriously taking care of the details in hardware; Lyra also proposes a set of synthesis and optimization techniques to automatically compile this "big-pipeline" program into multiple pieces of runnable chip-specific code that can be launched directly on the individual programmable switches of the target network. We built and evaluated Lyra. Lyra not only generates runnable real-world programs (in both P4 and NPL), but also uses up to 87.5% fewer hardware resources and up to 78% fewer lines of code than human-written programs. Ennan Zhai, Hongqiang Harry Liu, Rui Miao 0001, Yu Zhou 0008, Bingchuan Tian, Chen Sun 0005, Dennis Cai, Ming Zhang 0005, Minlan Yu |
SIGCOMM | 9 |
| 2020 | Accuracy, Scalability, Coverage: A Practical Configuration Verifier on a Global WANabstractThis paper presents Hoyan-- the first reported large scale deployment of configuration verification in a global-scale wide area network (WAN). Hoyan has been running in production for more than two years and is currently used for all critical configuration auditing and updates on the WAN. We highlight our innovative designs and real-life experience to make Hoyan accurate and scalable in practice. For accuracy under the inconsistencies of devices' vendor-specific behaviors (VSBs), Hoyan continuously discovers the flaws in device behavior models, thus aiding the operators in fixing the models. For scalability to verify our global WAN, Hoyan introduces a "global-simulation & local formal-modeling" strategy to model uncertainties in small scales and perform aggressive pruning of possibilities during the protocol simulations. Hoyan achieves near-100% verification accuracy after it detected and fixed O(10) VSBs on our WAN. Hoyan has prevented many potential service failures resulting from misconfiguration and reduced the failure rate of updates of our WAN by more than half in 2019. Fangdan Ye, Ennan Zhai, Hongqiang Harry Liu, Bingchuan Tian, Qiaobo Ye, Chunsheng Wang, Tianchen Guo, Duncheng She, Biao Cheng, Ming Zhang 0005, Rodrigo Fonseca |
SIGCOMM | 15 |
| 2020 | NFC+: Breaking NFC Networking Limits through Resonance EngineeringabstractCurrent UHF RFID systems suffer from two long-standing problems: 1) miss-reading non-line-of-sight or misoriented tags and 2) cross-reading undesired, distant tags due to multi-path reflections. This paper proposes a novel system, NFC+, to overcome the fundamental challenges. NFC+ is a magnetic field reader, which can inventory standard NFC tagged objects with a reasonably long range and arbitrary orientation. NFC+ achieves this by leveraging physical and algorithmic techniques based on magnetic resonance engineering. We build a prototype of NFC+ and conduct extensive evaluations in a logistic network. Comparing to UHF RFID, we find that NFC+ can reduce the miss-reading rate from 23% to 0.03%, and cross-reading rate from 42% to 0, for randomly oriented objects. NFC+ demonstrates high robustness for RFID unfriendly media (e.g., water bottles and metal cans). It can reliably read commercial NFC tags at a distance of up to 3 meters which, for the first time, enables NFC to be directly applied to practical logistics network applications. Renjie Zhao 0001, Purui Wang, Hongqiang Harry Liu, Xianshang Lin, Xinyu Zhang 0003, Chenren Xu, Ming Zhang 0005 |
SIGCOMM | 9 |
| 2019 | HPCC: high precision congestion controlabstractCongestion control (CC) is the key to achieving ultra-low latency, high bandwidth and network stability in high-speed networks. From years of experience operating large-scale and high-speed RDMA networks, we find the existing high-speed CC schemes have inherent limitations for reaching these goals. In this paper, we present HPCC (High Precision Congestion Control), a new high-speed CC mechanism which achieves the three goals simultaneously. HPCC leverages in-network telemetry (INT) to obtain precise link load information and controls traffic precisely. By addressing challenges such as delayed INT information during congestion and overreac-tion to INT information, HPCC can quickly converge to utilize free bandwidth while avoiding congestion, and can maintain near-zero in-network queues for ultra-low latency. HPCC is also fair and easy to deploy in hardware. We implement HPCC with commodity programmable NICs and switches. In our evaluation, compared to DCQCN and TIMELY, HPCC shortens flow completion times by up to 95%, causing little congestion even under large-scale incasts. Rui Miao 0001, Hongqiang Harry Liu, Lingbo Tang, Zheng Cao 0003, Ming Zhang 0005, Frank Kelly, Mohammad Alizadeh, Minlan Yu |
SIGCOMM | 8 |
| 2019 | Safely and automatically updating in-network ACL configurations with intent languageabstractIn-network Access Control List (ACL) is an important technique in ensuring network-wide connectivity and security. As cloud-scale WANs today constantly evolve in size and complexity, in-network ACL rules are becoming increasingly more complex. This presents a great challenge to the updating process of ACL configurations: network operators are frequently required to update "tangled" ACL rules across thousands of devices to meet diverse business requirements, and even a single ACL misconfiguration may lead to network disruptions. Such increasing challenges call for an automated system to improve the efficiency and correctness of ACL updates. This paper presents Jinjing, a system that aids Alibaba's network operators in automatically and correctly updating ACL configurations in Alibaba's global WAN. Jinjing allows the operators to express in a declarative language, named LAI, their update intent (e.g., ACL migration and traffic control). Then, Jinjing automatically synthesizes ACL update plans that satisfy their intent. At the heart of Jinjing, we develop a set of novel verification and synthesis techniques to rigorously guarantee the correctness of update plans. In Alibaba, our operators have used Jinjing to efficiently update their ACLs and have thus prevented significant service downtime. Bingchuan Tian, Xinyi Zhang 0003, Ennan Zhai, Hongqiang Harry Liu, Qiaobo Ye, Chunsheng Wang, Zhiming Ji, Yihong Sang, Ming Zhang 0005, Chen Tian 0001, Haitao Zheng 0001, Ben Y. Zhao |
SIGCOMM | 10 |
| 2017 | CherryPick: Adaptively Unearthing the Best Cloud Configurations for Big Data Analytics
Omid Alipourfard, Hongqiang Harry Liu, Jianshu Chen, Shivaram Venkataraman, Minlan Yu, Ming Zhang 0005 |
NSDI | 6 |
| 2017 | Guaranteeing Deadlines for Inter-Data Center TransfersabstractInter-data center wide area networks (inter-DC WANs) carry a significant amount of data transfers that require to be completed within certain time periods, or deadlines. However, very little work has been done to guarantee such deadlines. The crux is that the current inter-DC WAN lacks an interface for users to specify their transfer deadlines and a mechanism for provider to ensure the completion while maintaining high WAN utilization. In this paper, we address the problem by introducing a deadline-based network abstraction (DNA) for inter-DC WANs. DNA allows users to explicitly specify the amount of data to be delivered and the deadline by which it has to be completed. The malleability of DNA provides flexibility in resource allocation. Based on this, we develop a system calledAmoebathat implements DNA. Our simulations and test bed experiments show thatAmoeba, by harnessing DNA’s malleability, accommodates 15% more user requests with deadlines, while achieving 60% higher WAN utilization than prior solutions. Hong Zhang 0025, Kai Chen 0005, Wei Bai 0001, Dongsu Han, Chen Tian 0001, Hao Wang 0022, Haibing Guan, Ming Zhang 0005 |
IEEE/ACM Trans. Netw. | 8 |
| 2016 | Yoda: a highly available layer-7 load balancerabstractLayer-7 load balancing is a foundational building block of online services. The lack of offerings from major public cloud providers have left online services to build their own load balancers (LB), or use third-party LB design such as HAProxy. The key problem with such proxy-based design is each proxy instance is a single point of failure, as upon its failure, the TCP flow state for the connections with the client and server is lost which breaks the user flows. This significantly affects user experience and online services revenue. Rohan Gandhi, Y. Charlie Hu, Ming Zhang 0005 |
EuroSys | 3 |
| 2016 | Efficiently Delivering Online Services over Integrated Infrastructure
Hongqiang Harry Liu, Raajay Viswanathan, Matt Calder, Aditya Akella, Ratul Mahajan, Jitendra Padhye, Ming Zhang 0005 |
NSDI | 7 |
| 2015 | Guaranteeing deadlines for inter-datacenter transfersabstractInter-datacenter wide area networks (inter-DC WAN) carry a significant amount of data transfers that require to be completed within certain time periods, or deadlines. However, very little work has been done to guarantee such deadlines. The crux is that the current inter-DC WAN lacks an interface for users to specify their transfer deadlines and a mechanism for provider to ensure the completion while maintaining high WAN utilization. Hong Zhang 0025, Kai Chen 0005, Wei Bai 0001, Dongsu Han, Chen Tian 0001, Hao Wang 0022, Haibing Guan, Ming Zhang 0005 |
EuroSys | 8 |
| 2015 | Congestion Control for Large-Scale RDMA DeploymentsabstractModern datacenter applications demand high throughput (40Gbps) and ultra-low latency (< 10 μs per hop) from the network, with low CPU overhead. Standard TCP/IP stacks cannot meet these requirements, but Remote Direct Memory Access (RDMA) can. On IP-routed datacenter networks, RDMA is deployed using RoCEv2 protocol, which relies on Priority-based Flow Control (PFC) to enable a drop-free network. However, PFC can lead to poor application performance due to problems like head-of-line blocking and unfairness. To alleviates these problems, we introduce DCQCN, an end-to-end congestion control scheme for RoCEv2. To optimize DCQCN performance, we build a fluid model, and provide guidelines for tuning switch buffer thresholds, and other protocol parameters. Using a 3-tier Clos network testbed, we show that DCQCN dramatically improves throughput and fairness of RoCEv2 RDMA traffic. DCQCN is implemented in Mellanox NICs, and is being deployed in Microsoft's datacenters. Yibo Zhu 0001, Haggai Eran, Daniel Firestone, Chuanxiong Guo, Marina Lipshteyn, Yehonatan Liron, Jitendra Padhye, Shachar Raindel, Mohamad Haj Yahia, Ming Zhang 0005 |
SIGCOMM | 10 |
| 2015 | Packet-Level Telemetry in Large Datacenter NetworksabstractDebugging faults in complex networks often requires capturing and analyzing traffic at the packet level. In this task, datacenter networks (DCNs) present unique challenges with their scale, traffic volume, and diversity of faults. To troubleshoot faults in a timely manner, DCN administrators must a) identify affected packets inside large volume of traffic; b) track them across multiple network components; c) analyze traffic traces for fault patterns; and d) test or confirm potential causes. To our knowledge, no tool today can achieve both the specificity and scale required for this task. Yibo Zhu 0001, Nanxi Kang, Jiaxin Cao, Albert G. Greenberg, Guohan Lu, Ratul Mahajan, David A. Maltz, Ming Zhang 0005, Ben Y. Zhao, Haitao Zheng 0001 |
SIGCOMM | 9 |
| 2015 | Rubik: Unlocking the Power of Locality and End-point Flexibility in Cloud Scale Load Balancing
Rohan Gandhi, Y. Charlie Hu, Cheng-Kok Koh, Hongqiang Harry Liu, Ming Zhang 0005 |
USENIX ATC | 5 |
| 2014 | Duet: cloud scale load balancing with hardware and softwareabstractLoad balancing is a foundational function of datacenter infrastructures and is critical to the performance of online services hosted in datacenters. As the demand for cloud services grows, expensive and hard-to-scale dedicated hardware load balancers are being replaced with software load balancers that scale using a distributed data plane that runs on commodity servers. Software load balancers offer low cost, high availability and high flexibility, but suffer high latency and low capacity per load balancer, making them less than ideal for applications that demand either high throughput, or low latency or both. In this paper, we present Duet, which offers all the benefits of software load balancer, along with low latency and high availability -- at next to no cost. We do this by exploiting a hitherto overlooked resource in the data center networks -- the switches themselves. We show how to embed the load balancing functionality into existing hardware switches, thereby achieving organic scalability at no extra cost. For flexibility and high availability, Duet seamlessly integrates the switch-based load balancer with a small deployment of software load balancer. We enumerate and solve several architectural and algorithmic challenges involved in building such a hybrid load balancer. We evaluate Duet using a prototype implementation, as well as extensive simulations driven by traces from our production data centers. Our evaluation shows that Duet provides 10x more capacity than a software load balancer, at a fraction of a cost, while reducing latency by a factor of 10 or more, and is able to quickly adapt to network dynamics including failures. Rohan Gandhi, Hongqiang Harry Liu, Y. Charlie Hu, Guohan Lu, Jitendra Padhye, Ming Zhang 0005 |
SIGCOMM | 7 |
| 2014 | Dynamic scheduling of network updatesabstractWe present Dionysus, a system for fast, consistent network updates in software-defined networks. Dionysus encodes as a graph the consistency-related dependencies among updates at individual switches, and it then dynamically schedules these updates based on runtime differences in the update speeds of different switches. This dynamic scheduling is the key to its speed; prior update methods are slow because they pre-determine a schedule, which does not adapt to runtime conditions. Testbed experiments and data-driven simulations show that Dionysus improves the median update speed by 53--88% in both wide area and data center networks compared to prior methods. Xin Jin 0008, Hongqiang Harry Liu, Rohan Gandhi, Srikanth Kandula, Ratul Mahajan, Ming Zhang 0005, Jennifer Rexford, Roger Wattenhofer |
SIGCOMM | 6 |
| 2014 | Traffic engineering with forward fault correctionabstractFaults such as link failures and high switch configuration delays can cause heavy congestion and packet loss. Because it takes time to detect and react to faults, these conditions can last long---even tens of seconds. We propose forward fault correction (FFC), a proactive approach to handling faults. FFC spreads network traffic such that freedom from congestion is guaranteed under arbitrary combinations of up to k faults. We show how FFC can be practically realized by compactly encoding the constraints that arise from this large number of possible faults and solving them efficiently using sorting networks. Experiments with data from real networks show that, with negligible loss in overall network throughput, FFC can reduce data loss by a factor of 7--130 in well-provisioned networks, and reduce the loss of high-priority traffic to almost zero in well-utilized networks. Hongqiang Harry Liu, Srikanth Kandula, Ratul Mahajan, Ming Zhang 0005, David Gelernter |
SIGCOMM | 4 |
| 2014 | A network-state management serviceabstractWe present Statesman, a network-state management service that allows multiple network management applications to operate independently, while maintaining network-wide safety and performance invariants. Network state captures various aspects of the network such as which links are alive and how switches are forwarding traffic. Statesman uses three views of the network state. In observed state, it maintains an up-to-date view of the actual network state. Applications read this state and propose state changes based on their individual goals. Using a model of dependencies among state variables, Statesman merges these proposed states into a target state that is guaranteed to maintain the safety and performance invariants. It then updates the network to the target state. Statesman has been deployed in ten Microsoft Azure datacenters for several months, and three distinct applications have been built on it. We use the experience from this deployment to demonstrate how Statesman enables each application to meet its goals, while maintaining network-wide invariants. Ratul Mahajan, Jennifer Rexford, Ming Zhang 0005, Ahsan Arefin |
SIGCOMM | 5 |
| 2013 | Achieving high utilization with software-driven WANabstractWe present SWAN, a system that boosts the utilization of inter-datacenter networks by centrally controlling when and how much traffic each service sends and frequently re-configuring the network's data plane to match current traffic demand. But done simplistically, these re-configurations can also cause severe, transient congestion because different switches may apply updates at different times. We develop a novel technique that leverages a small amount of scratch capacity on links to apply updates in a provably congestion-free manner, without making any assumptions about the order and timing of updates at individual switches. Further, to scale to large networks in the face of limited forwarding table capacity, SWAN greedily selects a small set of entries that can best satisfy current demand. It updates this set without disrupting traffic by leveraging a small amount of scratch capacity in forwarding tables. Experiments using a testbed prototype and data-driven simulations of two production networks show that SWAN carries 60% more traffic than the current practice. Chi-Yao Hong, Srikanth Kandula, Ratul Mahajan, Ming Zhang 0005, Vijay Gill, Mohan Nanduri, Roger Wattenhofer |
SIGCOMM | 4 |
| 2013 | zUpdate: updating data center networks with zero lossabstractDatacenter networks (DCNs) are constantly evolving due to various updates such as switch upgrades and VM migrations. Each update must be carefully planned and executed in order to avoid disrupting many of the mission-critical, interactive applications hosted in DCNs. The key challenge arises from the inherent difficulty in synchronizing the changes to many devices, which may result in unforeseen transient link load spikes or even congestions. We present one primitive, zUpdate, to perform congestion-free network updates under asynchronous switch and traffic matrix changes. We formulate the update problem using a network model and apply our model to a variety of representative update scenarios in DCNs. We develop novel techniques to handle several practical challenges in realizing zUpdate as well as implement the zUpdate prototype on OpenFlow switches and deploy it on a testbed that resembles real DCN topology. Our results, from both real-world experiments and large-scale trace-driven simulations, show that zUpdate can effectively perform congestion-free updates in production DCNs. Hongqiang Harry Liu, Ming Zhang 0005, Roger Wattenhofer, David A. Maltz |
SIGCOMM | 3 |
| 2013 | Rake: Semantics Assisted Network-Based Tracing FrameworkabstractThe ability to trace request execution paths is critical for diagnosing performance faults in large-scale distributed systems. Previous black-box and white-box approaches are either inaccurate or invasive. We present a novel semantics-assisted gray-box tracing approach, called Rake, which can accurately trace individual request by observing network traffic. Rake infers the causality between messages by identifying polymorphic IDs in messages according to application semantics. To make Rake universally applicable, we design a Rake language so that users can easily describe necessary semantics of their applications while reusing the core Rake component. We evaluate Rake using a few popular distributed applications, including web search, distributed computing cluster, content provider network, and online chatting. Our results demonstrate Rake is much more accurate than the black-box approaches while requiring no modification to OS/applications. In the CoralCDN (a content distributed network) experiments, Rake links messages with much higher accuracy than WAP5, a state-of-the-art black-box approach. In the Hadoop (a distributed computing cluster platform) experiments, Rake helps reveal several previously unknown issues that may lead to performance degradation, including a RPC (Remote Procedure Call) abusing problem. Yao Zhao 0003, Yinzhi Cao, Yan Chen 0004, Ming Zhang 0005, Anup Goyal |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2012 | Where is the energy spent inside my app?: fine grained energy accounting on smartphones with EprofabstractWhere is the energy spent inside my app? Despite the immense popularity of smartphones and the fact that energy is the most crucial aspect in smartphone programming, the answer to the above question remains elusive. This paper first presents eprof, the first fine-grained energy profiler for smartphone apps. Compared to profiling the runtime of applications running on conventional computers, profiling energy consumption of applications running on smartphones faces a unique challenge, asynchronous power behavior, where the effect on a component's power state due to a program entity lasts beyond the end of that program entity. We present the design, implementation and evaluation of eprof on two mobile OSes, Android and Windows Mobile. Abhinav Pathak, Y. Charlie Hu, Ming Zhang 0005 |
EuroSys | 3 |
| 2012 | You Can Run, but You Can't Hide: Exposing Network Location for Targeted DoS Attacks in Cellular Networks
Zhiyun Qian, Zhaoguang Wang, Z. Morley Mao, Ming Zhang 0005, Yi-Min Wang |
NDSS | 5 |
| 2012 | NetPilot: automating datacenter network failure mitigationabstractDriven by the soaring demands for always-on and fast-response online services, modern datacenter networks have recently undergone tremendous growth. These networks often rely on commodity hardware to reach immense scale while keeping capital expenses under check. The downside is that commodity devices are prone to failures, raising a formidable challenge for network operators to promptly handle these failures with minimal disruptions to the hosted services. Daniel Turner, Chao-Chih Chen, David A. Maltz, Xiaowei Yang 0001, Ming Zhang 0005 |
SIGCOMM | 7 |
| 2011 | MicroTE: fine grained traffic engineering for data centersabstractThe effects of data center traffic characteristics on data center traffic engineering is not well understood. In particular, it is unclear how existing traffic engineering techniques perform under various traffic patterns, namely how do the computed routes differ from the optimal routes. Our study reveals that existing traffic engineering techniques perform 15% to 20% worse than the optimal solution. We find that these techniques suffer mainly due to their inability to utilize global knowledge about flow characteristics and make coordinated decision for scheduling flows. Theophilus Benson, Ashok Anand, Aditya Akella, Ming Zhang 0005 |
CoNEXT | 4 |
| 2011 | Fine-grained power modeling for smartphones using system call tracingabstractAccurate, fine-grained online energy estimation and accounting of mobile devices such as smartphones is of critical importance to understanding and debugging the energy consumption of mobile applications. We observe that state-of-the-art, utilization-based power modeling correlates the (actual) utilization of a hardware component with its power state, and hence is insufficient in capturing several power behavior not directly related to the component utilization in modern smartphones. Such behavior arise due to various low level power optimizations programmed in the device drivers. We propose a new, system-call-based power modeling approach which gracefully encompasses both utilization-based and non-utilization-based power behavior. We present the detailed design of such a power modeling scheme and its implementation on Android and Windows Mobile. Our experimental results using a diverse set of applications confirm that the new model significantly improves the fine-grained as well as whole-application energy consumption accuracy. We further demonstrate fine-grained energy accounting enabled by such a fined-grained power model, via amanually implemented eprof, the energy counterpart of the classic gprof tool, for profiling application energy drain. Abhinav Pathak, Y. Charlie Hu, Ming Zhang 0005, Paramvir Bahl, Yi-Min Wang |
EuroSys | 3 |
| 2011 | Bootstrapping energy debugging on smartphones: a first look at energy bugs in mobile devicesabstractThis paper argues that a new class of bugs faced by millions of smartphones, energy bugs or ebugs, have become increasingly prominent that already they have led to significant user frustrations. We take a first look at this emerging important technical challenge faced by the smartphones, ebugs, broadly defined as an error in the system (application, OS, hardware, firmware, external conditions or combination) that causes an unexpected amount of high energy consumption by the system as a whole. We first present a taxonomy of the kinds of ebugs based on mining over 39K posts (1.2M before filtering) from 4 online mobile user forum and mobile OS bug repositories. The taxonomy shows the highly diverse nature of smartphone ebugs. We then propose a roadmap towards developing a systematic diagnosing framework for debugging ebugs on smartphones. Abhinav Pathak, Y. Charlie Hu, Ming Zhang 0005 |
HotNets | 3 |
| 2011 | Latency inflation with MPLS-based traffic engineeringabstractWhile MPLS has been extensively deployed in recent years, little is known about its behavior in practice. We examine the performance of MPLS in Microsoft's online service network (MSN), a well-provisioned multi-continent production network connecting tens of data centers. Using detailed traces collected over a 2-month period, we find that many paths experience significantly inflated latencies. We correlate occurrences of latency inflation with routers, links, and DC-pairs. This analysis sheds light on the causes of latency inflation and suggests several avenues for alleviating the problem. Abhinav Pathak, Ming Zhang 0005, Y. Charlie Hu, Ratul Mahajan, David A. Maltz |
Internet Measurement Conference | 2 |
| 2011 | Rake: Semantics assisted network-based tracing frameworkabstractThe ability to trace request execution paths is critical for diagnosing performance faults in large-scale distributed systems. Previous black-box and white-box approaches are either inaccurate or invasive. We present a novel semantics-assisted gray-box tracing approach, called Rake, which can accurately trace individual request by observing network traffic. Rake infers the causality between messages by identifying polymorphic IDs in messages according to application semantics. To make Rake universally applicable, we design a Rake language so that users can easily describe necessary semantics of their applications while reusing the core Rake component. We evaluate Rake using a few popular distributed applications, including web search, distributed computing cluster, content provider network, and online chatting. Our results demonstrate Rake is much more accurate than the black-box approaches while requiring no modification to OS/applications. In the CoralCDN (a content distributed network) experiments, Rake links messages with much higher accuracy than WAP5, a state-of-the-art black-box approach. In the Hadoop (a distributed computing cluster platform) experiments, Rake helps reveal several previously unknown issues that may lead to performance degradation, including a RPC (Remote Procedure Call) abusing problem. Yao Zhao 0003, Yinzhi Cao, Yan Chen 0004, Ming Zhang 0005, Anup Goyal |
IWQoS | 4 |
| 2011 | Switchboard: a matchmaking system for multiplayer mobile gamesabstractSupporting interactive, multiplayer games on mobile phones over cellular networks is a difficult problem. It is particularly relevant now with the explosion of mostly single-player or turn-based games on mobile phones. The challenges stem from the highly variable performance of cellular networks and the need for scalability (not burdening the cellular infrastructure, nor any server resources that a game developer deploys). We have built a service for matchmaking in mobile games -- assigning players to games such that game settings are satisfied as well as latency requirements for an enjoyable game. This requires solving two problems. First, the service needs to know the cellular network latency between game players. Second, the service needs to quickly group players into viable game sessions. In this paper, we present the design of our service, results from our experiments on predicting cellular latency, and results from efficiently grouping players into games. Justin Manweiler, Sharad Agarwal, Ming Zhang 0005, Romit Roy Choudhury, Paramvir Bahl |
MobiSys | 3 |
| 2011 | CloudProphet: towards application performance prediction in cloudabstractChoosing the best-performing cloud for one's application is a critical problem for potential cloud customers. We propose CloudProphet, a trace-and-replay tool to predict a legacy application's performance if migrated to a cloud infrastructure. CloudProphet traces the workload of the application when running locally, and replays the same workload in the cloud for prediction. We discuss two key technical challenges in designing CloudProphet, and some preliminary results using a prototype implementation. Ang Li 0002, Xuanran Zong, Srikanth Kandula, Xiaowei Yang 0001, Ming Zhang 0005 |
SIGCOMM | 5 |
| 2011 | An untold story of middleboxes in cellular networksabstractThe use of cellular data networks is increasingly popular as network coverage becomes more ubiquitous and many diverse user-contributed mobile applications become available. The growing cellular traffic demand means that cellular network carriers are facing greater challenges to provide users with good network performance and energy efficiency, while protecting networks from potential attacks. To better utilize their limited network resources while securing the network and protecting client devices the carriers have already deployed various network policies that influence traffic behavior. Today, these policies are mostly opaque, though they directly impact application designs and may even introduce network vulnerabilities. Zhaoguang Wang, Zhiyun Qian, Z. Morley Mao, Ming Zhang 0005 |
SIGCOMM | 5 |
| 2010 | CloudCmp: comparing public cloud providersabstractWhile many public cloud providers offer pay-as-you-go computing, their varying approaches to infrastructure, virtualization, and software services lead to a problem of plenty. To help customers pick a cloud that fits their needs, we develop CloudCmp, a systematic comparator of the performance and cost of cloud providers. CloudCmp measures the elastic computing, persistent storage, and networking services offered by a cloud along metrics that directly reflect their impact on the performance of customer applications. CloudCmp strives to ensure fairness, representativeness, and compliance of these measurements while limiting measurement cost. Applying CloudCmp to four cloud providers that together account for most of the cloud customers today, we find that their offered services vary widely in performance and costs, underscoring the need for thoughtful provider selection. From case studies on three representative cloud applications, we show that CloudCmp can guide customers in selecting the best-performing provider for their applications. Ang Li 0002, Xiaowei Yang 0001, Srikanth Kandula, Ming Zhang 0005 |
Internet Measurement Conference | 4 |
| 2010 | Anatomizing application performance differences on smartphonesabstractThe use of cellular data networks is increasingly popular due to the widespread deployment of 3G technologies and the rapid adoption of smartphones, such as iPhone and GPhone. Besides email and web browsing, a variety of network applications are now available, rendering smartphones potentially useful substitutes for their desktop counterparts. Nevertheless, the performance of smartphone applications in the wild is still poorly understood due to a lack of systematic measurement methodology. Junxian Huang 0001, Birjodh Singh Tiwana, Z. Morley Mao, Ming Zhang 0005, Paramvir Bahl |
MobiSys | 5 |
| 2010 | WebProphet: Automating Performance Prediction for Web Services
Zhichun Li, Ming Zhang 0005, Zhaosheng Zhu, Yan Chen 0004, Albert G. Greenberg, Yi-Min Wang |
NSDI | 2 |
| 2010 | Optimizing Cost and Performance in Online Service Provider Networks
Zheng Zhang 0009, Ming Zhang 0005, Albert G. Greenberg, Y. Charlie Hu, Ratul Mahajan, Blaine Christian |
NSDI | 2 |
| 2009 | Detecting traffic differentiation in backbone ISPs with NetPoliceabstractTraffic differentiations are known to be found at the edge of the Internet in broadband ISPs and wireless carriers [13, 2]. The ability to detect traffic differentiations is essential for customers to develop effective strategies for improving their application performance. We build a system, called NetPolice, that enables detection of content- and routing-based differentiations in backbone ISPs. NetPolice is easy to deploy since it only relies on loss measurement launched from end hosts. The key challenges in building NetPolice include selecting an appropriate set of probing destinations and ensuring the robustness of detection results to measurement noise. Ying Zhang 0022, Z. Morley Mao, Ming Zhang 0005 |
Internet Measurement Conference | 3 |
| 2008 | Ascertaining the Reality of Network Neutrality Violation in Backbone ISPs
Ying Zhang 0022, Z. Morley Mao, Ming Zhang 0005 |
HotNets | 3 |
| 2008 | Uncovering Performance Differences Among Backbone ISPs with Netdiff
Ratul Mahajan, Ming Zhang 0005, Lindsey Poole, Vivek S. Pai |
NSDI | 2 |
| 2008 | Effective Diagnosis of Routing Disruptions from End Systems
Ying Zhang 0022, Z. Morley Mao, Ming Zhang 0005 |
NSDI | 3 |
| 2008 | Automating Network Application Dependency Discovery: Experiences, Limitations, and New Solutions
Xu Chen 0028, Ming Zhang 0005, Z. Morley Mao, Paramvir Bahl |
OSDI | 2 |
| 2007 | Towards highly reliable enterprise network services via inference of multi-level dependenciesabstractLocalizing the sources of performance problems in large enterprise networks is extremely challenging. Dependencies are numerous, complex and inherently multi-level, spanning hardware and software components across the network and the computing infrastructure. To exploit these dependencies for fast, accurate problem localization, we introduce an Inference Graph model, which is well-adapted to user-perceptible problems rooted in conditions giving rise to both partial service degradation and hard faults. Further, we introduce the Sherlock system to discover Inference Graphs in the operational enterprise, infer critical attributes, and then leverage the result to automatically detect and localize problems. To illuminate strengths and limitations of the approach, we provide results from a prototype deployment in a large enterprise network, as well as from testbed emulations and simulations. In particular, we find that taking into account multi-level structure leads to a 30% improvement in fault localization, as compared to two-level approaches. Paramvir Bahl, Ranveer Chandra, Albert G. Greenberg, Srikanth Kandula, David A. Maltz, Ming Zhang 0005 |
SIGCOMM | 6 |
| 2006 | Discovering Dependencies for Network Management
Paramvir Bahl, Paul Barham 0001, Richard Black, Ranveer Chandra, Moisés Goldszmidt, Rebecca Isaacs, Srikanth Kandula, John MacCormick, David A. Maltz, Richard Mortier, Michal Wawrzoniak, Ming Zhang 0005 |
HotNets | 13 |
| 2006 | WiFiProfiler: cooperative diagnosis in wireless LANsabstractWhile 802.11-based wireless hotspots are proliferating, users often have little recourse when the network does not work or performs poorly for them. They are left trying to manually debug the problem, which can be a frustrating and disruptive process. The users' troubles are compounded by the absence of network administrators or an IT department to turn to in many 802.11 hotspot settings (e.g., cafes, airports, conferences).We present WiFiProfiler, a system in which wireless hosts cooperate to diagnose and possibly resolve network problems in an automated manner, without requiring any infrastructural support. The key observation is that even if a host's wireless link to an access point is not working, the host is often within the range of other wireless nodes and is in a position to communicate with them (a little) peer-to-peer. We leverage this ability to create a shared information plane, which enables wireless hosts to exchange a range of information about their network settings and the health of their network connectivity. By aggregating and correlating such information across multiple wireless hosts, we infer the likely cause of the problem. Our implementation on Windows XP shows that WiFiProfiler is effective in diagnosing a range of problems and imposes a low overhead on the participating hosts. Ranveer Chandra, Venkat N. Padmanabhan, Ming Zhang 0005 |
MobiSys | 3 |
| 2006 | How DNS Misnaming Distorts Internet Topology Mapping
Ming Zhang 0005, Yaoping Ruan, Vivek S. Pai, Jennifer Rexford |
USENIX ATC, General Track | 1 |
| 2004 | Network-Embedded Programmable Storage and Its Applications
Sumeet Sobti, Junwen Lai, Yilei Shao, Chi Zhang 0070, Ming Zhang 0005, Fengzhou Zheng, Arvind Krishnamurthy, Randolph Y. Wang |
NETWORKING | 6 |
| 2004 | PlanetSeer: Internet Path Failure Monitoring and Characterization in Wide-Area Services
Ming Zhang 0005, Chi Zhang 0070, Vivek S. Pai, Larry L. Peterson, Randolph Y. Wang |
OSDI | 1 |
| 2004 | A Transport Layer Approach for Improving End-to-End Performance and Robustness Using Redundant Paths
Ming Zhang 0005, Junwen Lai, Arvind Krishnamurthy, Larry L. Peterson, Randolph Y. Wang |
USENIX ATC, General Track | 1 |
| 2003 | RR-TCP: A Reordering-Robust TCP with DSACKabstractTCP performs poorly on paths that reorder packets significantly, where it misinterprets out-of-order delivery as packet loss. The sender responds with a fast retransmit though no actual loss has occurred. These repeated false fast retransmits keep the sender's window small, and severely degrade the throughput it attains. Requiring nearly in-order delivery needlessly restricts and complicates Internet routing systems and routers. Such beneficial systems as multi-path routing and parallel packet switches are difficult to deploy in a way that preserves ordering. Toward a more reordering-tolerant Internet architecture, we present enhancements to TCP that improve the protocol's robustness to reordered and delayed packets. We extend the sender to detect and recover from false fast retransmits using DSACK information, and to avoid false fast retransmits proactively, by adaptively varying dupthresh. Our algorithm is the first that adaptively balances increasing dupthresh, to avoid false fast retransmits, and limiting the growth of dupthresh, to avoid unnecessary timeouts. Finally, we demonstrate that TCP's RTO estimator tolerates delayed packets poorly, and present enhancements to it that ensure it is sufficiently conservative, without using timestamps or additional TCP header hits. Our simulations show that these enhancements significantly improve TCP's performance over paths that reorder or delay packets. Ming Zhang 0005, Brad Karp, Sally Floyd, Larry L. Peterson |
ICNP | 1 |
| 2002 | Probabilistic Packet Scheduling: Achieving Proportional Share Bandwidth Allocation for TCP FlowsabstractThis paper describes and evaluates a probabilistic packet scheduling (PPS) algorithm for providing different levels of service to TCP flows. With our approach, each router defines a local currency in terms of tickets and assigns tickets to its inputs based on contractual agreements with its upstream routers. A flow is tagged with tickets to represent the relative share of bandwidth it should receive at each link. When multiple flows share the same bottleneck, the bandwidth that each flow obtains is proportional to the relative tickets assigned to that flow. Simulations show that PPS does a better job of proportionally allocating bandwidth than DiffServ and weighted CSFQ. In addition, PPS accommodates flows that cross multiple currency domains. Ming Zhang 0005, Randolph Y. Wang, Larry L. Peterson, Arvind Krishnamurthy |
INFOCOM | 1 |