VLDB 2026 Research / reviewers in the wild / expert
Ying Zhang 0022
dblp:13/6769-22
· DBLP profile ↗
89ranked-venue papers
18as first author
42since 2021 · last 2026
0000-0003-2736-5694ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 69 · 13 first-author · 37 since 2021Systems, architecture and hardware · 12 · 4 first-author · 3 since 2021Security and privacy · 7 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Matryoshka: Realizing Hyperscale Data Center Network Design for the AI Era
Yan Cai 0018, Jialong Li 0006, Kutalmis Akpinar, Hany Morsy, Sunil Khaunte, Yiting Xia, Ying Zhang 0022 |
NSDI | 9 |
| 2026 | Express Lane to Efficiency and Reliability: Multi-Dimensional Control in Meta's Express Backbone Network
Faisal Iqbal, Vitaly Neganov, Brian Chang, Rong Rong, Yuanjun Yao, Marek Denis, Alexandru Manea, Anton Marchenko, Ulas Kozat, Aditya Akella, Ying Zhang 0022 |
NSDI | 12 |
| 2026 | Phantora: Maximizing Code Reuse in Simulation-based Machine Learning System Performance Estimation
Jianxing Qin, Jingrong Chen 0002, Xinhao Kong, Tianjun Yuan, Zhaodong Wang, Ying Zhang 0022, Tingjun Chen, Alvin R. Lebeck, Danyang Zhuo |
NSDI | 8 |
| 2026 | Enabling AI Network Cross-Layer Design and Operations with Arcadia: A Simulation Platform at Scale
Zhaodong Wang, Satyajeet Ahuja, Mohammad Noormohammadpour, Gregory R. Steinbrecher, Thomas Fuller, Kevin Quirk, Mikel Jimenez Fernandez, Abhinav Triguna, Yan Cai 0018, Steve Politis, Petr Lapukhov, Naader Hasani, Ying Zhang 0022 |
NSDI | 15 |
| 2026 | Planogram: A Multi-dimensional Physical Location Planning System for DC NetworksabstractMeta's data centers underpin a vast array of Internet services and have faced unprecedented demand due to the rapid expansion of AI workloads. The traditional approach of building standardized data centers is increasingly challenged by the exponential growth in required capacity that is now sourced in a variety of non-standard physical environments and data center designs. This shift introduces a complex challenge: how to rapidly and repeatably design custom data center networks that balance multiple, often conflicting, objectives across diverse engineering disciplines. Richard Cziva, Alexander Mafusalov, Shrinivas Petale, Abhinav Triguna, Manikantan Kr, Susana Contrera, Jimmy Williams, Alexey Andreyev, Tian Fang, Satyajeet Ahuja, Ying Zhang 0022 |
SIGCOMM | 13 |
| 2026 | Achieving Network Efficiency Through Service Collaborative Capacity Sharing and EnforcementabstractMeta's rapid expansion in users, business operations, and AI workloads is straining our backbone network, while physical constraints—such as fiber, space, and power—limit the speed of capacity growth. To address these challenges, we present a service-aware network capacity planning suite that systematically improves network efficiency with a service collaboration approach. We propose the "safe capacity" abstraction which enables services to incorporate current and projected network conditions into their compute and storage allocation decisions. We introduce a hose-carving method that efficiently translates service-level traffic demands into detailed traffic matrices, allowing for more precise bandwidth allocation. To promote responsible network usage, we design a network rate card which attributes network consumption to individual services, incentivizing optimization and resource trade-offs. Additionally, new enforcement features at the end-host layer dynamically adjust resource allocations and traffic flows at runtime to maximize utilization. This paper is the first to detail a collaborative, service-aware approach to backbone network efficiency at Meta scale. Based on years of operational experience, we share practical insights and highlight new directions for research in network efficiency. Vinayak Dangui, Alaleh Razmjoo, Guanqing Yan, Mahesh Nayak, Mansi Babbar, Brian Bierig, Tejas Birajdar, Prabhakaran Ganesan, Lilian Liu, Matt Maia, Jerry Yang, Shrinivas Petale, Satyajeet Ahuja, Abhinav Triguna, Guyue Liu, Ying Zhang 0022 |
SIGCOMM | 16 |
| 2026 | CacheFlare: Optimizing Cold Content Performance in CDNsabstractExisting research on Content Delivery Networks (CDNs) predominantly focuses on optimizing the delivery of hot content—popular items that attract frequent and repeated access from large user bases. However, at Meta, we have identified that cold content, which is less popular and accessed infrequently, poses significant challenges for user experience, especially with large volumes of direct messaging media. Through extensive analysis of Meta's CDN data, we quantify the impact of cold content performance on user experience and reveal substantial differences across geographic regions. Motivated by these findings, we propose CacheFlare, a suite of production-deployed solutions designed to enhance Quality of Experience (QoE) by increasing CDN hit rates for cold content. Through experiments and production validation, we show that even cold content can benefit from improved caching strategies, yielding up to a 10% improvement across a range of network performance metrics. Tiansheng Zhang, Ahmed Kamal, Jianfeng Tang, Huapeng Zhou, Thilan Ganegedara, Sanjay Sane, Ben Vallis, Theophilus Benson, Ying Zhang 0022 |
SIGCOMM | 12 |
| 2025 | One to Many: Closing the Bandwidth Gap in AI Datacenters with Scalable MulticastabstractAI training now floods datacenter fabrics with thousands of simultaneous collectives, yet most frameworks still move data the hard way: O(N) unicasts for a group of N processors. Classic multicast could slash those bytes but has long been deemed unscalable: computing an optimal tree in an asymmetric Clos is NP-hard and group-specific rules quickly exhaust switch TCAM. Sepehr Abdous, Jinqi Lu, Jiacheng Wan, Erfan Sharafzadeh, Ying Zhang 0022, Soudeh Ghorbani |
HotNets | 5 |
| 2025 | Congestion Patterns in a Large-scale RDMA DatacenterabstractRDMA datacenters are proliferating to meet the demand of emerging workloads such as AI training and inference as well as distributed storage. This trend has opened up a critical knowledge gap: the traffic characteristics of congestion in these networks remain unknown. We do not know, for example, which layers of the network are the most congested, if the network is load balanced effectively, how long congestion events last, and how accurate existing telemetry systems are in capturing congestion. This paper bridges this gap by investigating congestion in a large-scale RDMA datacenter dedicated to distributed AI training. We provide insights into three specific congestion patterns: (a) location and distribution in the network, (b) burstiness, e.g., the duration and synchrony of bursts, and (c) observability using existing telemetry methods. We show, for instance, that the deployment of Priority Flow Control (PFC) in RDMA networks has shifted the location of congestion one level up: from the edge-host in legacy TCP/IP datacenters to the network core in RDMA datacenters. At the same time, we show that the same protocol enables us to observe and understand congestion better, even bursty events. The findings of this research reveal open challenges for measuring, characterizing, and managing congestion in RDMA networks, paving the way for future research. Soudeh Ghorbani, Yimeng Zhao, Srikanth Sundaresan, Ying Zhang 0022, Yijing Zeng, Abhigyan Sharma, Prashanth Kannan, Cristian Lumezanu |
IMC | 4 |
| 2025 | Heimdall: Towards Risk-Aware Network Management Outsourcing
Yuejie Wang, Qiutong Men, Yongting Chen, Jiajin Liu, Gengyu Chen, Ying Zhang 0022, Guyue Liu, Vyas Sekar |
NDSS | 6 |
| 2025 | Enabling Efficient GPU Communication over Multiple NICs with FuseLink
Zhenghang Ren, Zilong Wang 0007, Wenxue Li 0004, Kaiqiang Xu, Xudong Liao, Yijun Sun, Bowen Liu 0002, Han Tian, Junxue Zhang 0001, Mingfei Wang, Zhizhen Zhong, Guyue Liu, Ying Zhang 0022, Kai Chen 0005 |
OSDI | 15 |
| 2025 | Hattrick: Solving Multi-Class TE using Neural ModelsabstractWhile recent work shows ML-based approaches are a promising alternative to conventional optimization methods for Traffic Engineering (TE), existing research is limited to a single traffic class. In this paper, we present Hattrick, the first ML-based approach for handling multiple traffic classes, a key requirement of cloud and ISP WANs. As part of Hattrick we have developed (i) a novel neural architecture aligned with the sequence of optimization problems in multiclass TE; and (ii) a variant of classical multitask learning methods to deal with the unique challenge of optimizing multiple metrics that have a precedence relationship. Evaluations on a large private WAN and other public datasets show Hattrick outperforms state-of-the-art optimization-based multiclass TE methods by better coping with prediction error - e.g., for GEANT, Hattrick outperforms SWAN by 5.48% to 19.3% across classes when considering the traffic that can be supported 99% of the time. Abd AlRhman AlQiam, Zhuocong Li, Satyajeet Ahuja, Zhaodong Wang, Ying Zhang 0022, Sanjay G. Rao, Bruno Ribeiro 0001, Mohit Tawarmalani |
SIGCOMM | 5 |
| 2025 | MixNet: A Runtime Reconfigurable Optical-Electrical Fabric for Distributed Mixture-of-Experts TrainingabstractMixture-of-Expert (MoE) models outperform conventional models by selectively activating different subnets, named experts, on a per-token basis. This gated computation generates dynamic communications that cannot be determined beforehand, challenging the existing GPU interconnects that remain static during distributed training. In this paper, we advocate for a first-of-its-kind system, called MixNet, that unlocks topology reconfiguration during distributed MoE training. Towards this vision, we first perform a production measurement study and show that the MoE dynamic communication pattern has strong locality, alleviating the need for global reconfiguration. Based on this, we design and implement a regionally reconfigurable high-bandwidth domain that augments existing electrical interconnects using optical circuit switching (OCS), achieving scalability while maintaining rapid adaptability. We build a fully functional MixNet prototype with commodity hardware and a customized collective communication runtime. Our prototype trains state-of-the-art MoE models with in-training topology reconfiguration across 32 A100 GPUs. Large-scale packet-level simulations show that MixNet achieves performance comparable to a non-blocking fat-tree fabric while boosting the networking cost efficiency (e.g., performance per dollar) of four representative MoE models by 1.2×–1.5× and 1.9×–2.3× at 100 Gbps and 400 Gbps link bandwidths, respectively. Xudong Liao, Yijun Sun, Han Tian, Xinchen Wan, Yilun Jin, Zilong Wang 0007, Zhenghang Ren, Wenxue Li 0004, Kin Fai Tse, Zhizhen Zhong, Guyue Liu, Ying Zhang 0022, Xiaofeng Ye, Yiming Zhang 0003, Kai Chen 0005 |
SIGCOMM | 13 |
| 2025 | Centralium: A Hybrid Route-Planning Framework for Large-Scale Data Center Network MigrationsabstractMeta's data center networks have relied on BGP for routing and interconnectivity due to its scalability and simplicity. However, as our network evolves, BGP's limitations in supporting complex network migrations have become apparent. These migrations require customized routing at each intermediate step, considering topological properties. In this paper, we highlight the unique challenges of production migration and illustrate how native BGP falls short, unable to encode both sequential and spatial conditions. To address this challenge, we introduce a novel Route Planning Abstraction (RPA) that augments the BGP protocol to support migration. It enables centralized route planning while maintaining distributed enforcement. We demonstrate the power of this abstraction by developing Centralium, a hybrid route-planning framework with a centralized controller, and over ten use cases to support various migrations in production. Our production experience with Centralium, deployed alongside BGP in large-scale data centers, has shown substantial reduction in the time and risk of network migration operations. Yikai Lin, Mohab Gawish, Shih-Hao Tseng, Lixin Gao 0001, Cen Zhao, John Tracey, Sunyi Shao, Hyojeong Kim, Ying Zhang 0022 |
SIGCOMM | 9 |
| 2025 | Unlocking Superior Performance in Reconfigurable Data Center Networks with Credit-Based TransportabstractThe large-scale, end-to-end implementation of microsecond-switched reconfigurable data center networks (RDCNs), coupled with innovative routing and topology designs that provide continuous routes abstracting away frequent topology changes, demonstrates promise as a viable alternative to Clos networks in the post-Moore's Law era. However, the gap remains in transport performance, with current transport solutions falling short of unlocking their full performance. In this paper, we introduce Flare, a novel credit-based transport protocol that ensures reliable traffic delivery, low latency, and leverages the rapidly reconfiguring circuits of the RDCN to opportunistically route traffic over short paths, maximizing throughput. In simulations, Flare enables RDCNs to outperform Clos networks, achieving up to 1.15× higher throughput even under adversarial traffic. Additionally, it delivers up to 2× and 1.5× higher throughput than NDP and ExpressPass, and up to 10×, 15×, and 3.5× shorter flow completion time (FCT) than ExpressPass, TDTCP, and Bolt. Our testbed implementation further demonstrates the feasibility of Flare's mechanisms with programmable switches and DPDK. Federico De Marchi 0002, Jialong Li 0006, Ying Zhang 0022, Wei Bai 0001, Yiting Xia |
SIGCOMM | 3 |
| 2025 | PreTE: Traffic Engineering with Predictive FailuresabstractFiber links in wide-area networks (WANs) are exposed to complicated environments and hence are vulnerable to failures like fiber cuts. The conventional approach of using static probabilistic failures falls short in fiber-cut scenarios because these fiber cuts are rare but disruptive, making it difficult for network operators to balance network utilization and availability in WAN traffic engineering. Our large-scale measurements of per-second optical-layer data reveal that the fiber's failure probability increases by several orders of magnitude when experiencing a rare and ephemeral degradation state. Therefore, we present a novel traffic engineering (TE) system called PreTE to factor in the dynamic fiber cut probabilities directly into TE systems. At the core of the PreTE system, fiber degradation facilitates failure predictions and traffic tunnels to be proactively updated, followed by traffic allocation optimizations among updated tunnels. We evaluate PreTE using a production-level WAN testbed and large-scale simulations. The testbed evaluation quantifies PreTE's runtime to demonstrate the feasibility to implement in large-scale WANs. Our large-scale simulation results show that PreTE can support up to 2× more demand at the same level of availability as compared to existing TE schemes. Congcong Miao, Zhizhen Zhong, Arpit Gupta, Ying Zhang 0022, Zekun He, Xianneng Zou, Jilong Wang 0001 |
SIGCOMM | 5 |
| 2025 | Intent-Driven Network Management with Multi-Agent LLMs: The Confucius FrameworkabstractAdvancements in Large Language Models (LLMs) are significantly transforming network management practices. In this paper, we present our experience developing Confucius, a multi-agent framework for network management at Meta. We model network management workflows as directed acyclic graphs (DAGs) to aid planning. Our framework integrates LLMs with existing management tools to achieve seamless operational integration, employs retrieval-augmented generation (RAG) to improve long-term memory, and establishes a set of primitives to systematically support human/model interaction. To ensure the accuracy of critical network operations, Confucius closely integrates with existing network validation methods and incorporates its own validation framework to prevent regressions. Remarkably, Confucius is a production-ready LLM development framework that has been operational for two years, with over 60 applications onboarded. To our knowledge, this is the first report on employing multi-agent LLMs for hyper-scale networks. Zhaodong Wang, Samuel Lin, Guanqing Yan, Soudeh Ghorbani, Minlan Yu, Jiawei Zhou 0012, Nathan Hu, Lopa Baruah, Sam Peters, Srikanth Kamath, Jerry Yang, Ying Zhang 0022 |
SIGCOMM | 12 |
| 2025 | An Anatomy of Token-Based Congestion ControlabstractCongestion control protocols play a vital role in enhancing the performance of various applications within datacenter networks. While reactive congestion control (RCC) protocols are widely deployed in commercial datacenters, the research community has actively explored token-based proactive congestion control (TCC) protocols to further push the boundaries of performance. However, despite the emergence of numerous TCC variants, there has been a lack of systematic exploration in the design space of TCC. This paper aims to bridge this gap by proposing a framework for understanding the design choices within the TCC approach. In this study, we systematically analyze different design choices of TCC approaches and leverage this understanding to develop a novel TCC protocol called ToCC. To implement ToCC, we address a set of challenges and deploy it in NP-based smart NICs. We compare ToCC with state-of-the-art TCC and RCC protocols through extensive large-scale simulations and testbed evaluations. The results demonstrate that ToCC exhibits robustness in achieving low latency across various scenarios. Additionally, ToCC effectively reduces buffer occupancy by 4.8 times compared to existing approaches, and under incast scenarios, it significantly shortens flow completion time by up to 90%. Congestion control protocols are crucial for optimizing the performance of datacenter network applications. Although reactive congestion control (RCC) protocols are commonly used in commercial datacenters, researchers have been exploring token-based proactive congestion control (TCC) protocols to further enhance network performance. Despite the development of numerous TCC variants, there has not been a thorough examination of the design space of TCC protocols until now. This paper aims to address this gap by introducing a framework for understanding the design choices within the TCC approach for TCC protocols. By analyzing various design aspects of TCC approaches, we create a novel TCC protocol called ToCC. At the central of ToCC design is that it leverages congestion control mechanisms over tokens. To implement ToCC, we tackle several challenges and integrate it into NP-based smart NICs. Comparing ToCC with state-of-the-art TCC and RCC protocols through extensive large-scale simulations and testbed evaluations, we find that ToCC consistently achieves low latency across different scenarios. Moreover, ToCC significantly reduces buffer occupancy by 4.8 times compared to existing methods, and during incast scenarios, it decreases flow completion time by up to 90%. Chang Liu 0001, Qingyue Wang, Lu Lu 0016, Xiaoliang Wang 0001, Fu Xiao 0001, Ying Zhang 0022, Wan-Chun Dou, Guihai Chen, Chen Tian 0001 |
IEEE Trans. Netw. | 8 |
| 2024 | Understanding Communication Characteristics of Distributed TrainingabstractCommunication is pivotal in distributed training and a thorough understanding of its characteristics is essential for future optimizations. However, prior works are limited, either focusing on customized optimizations or conducting incomplete explorations on communication characteristics. In this work, we systematically analyze the communication characteristics of distributed training, considering two key aspects of communication: pattern and overhead, and assessing a broad spectrum of determinant factors. In particular, we extensively investigate the features of communication patterns, such as predictability, and comprehensively evaluate the impact of various factors on communication overhead. Additionally, we develop and validate an analytical formulation to estimate communication overhead, providing a mathematical understanding of models with predictability. Wenxue Li 0004, Xiangzhou Liu, Yilun Jin, Han Tian, Zhizhen Zhong, Guyue Liu, Ying Zhang 0022, Kai Chen 0005 |
APNet | 8 |
| 2024 | Occam: A Programming System for Reliable Network ManagementabstractThe complexity of large networks makes their management a daunting task. State-of-the-art network management tools use workflow systems for automation, but they do not adequately address the substantial challenges in operation reliability. This paper presents Occam, a programming system that simplifies the development of reliable network management tasks. We leverage the fact that most modern network management systems are backed with a source-of-truth database, and thus customize database techniques to the context of network management. Occam exposes an easy-to-use programming model for network operators to express the key management logic, while shielding them from reliability concerns, such as operational conflicts and task atomicity. Instead, the Occam runtime provides these reliability guardrails automatically. Our evaluation demonstrates Occam's effectiveness in simplifying management tasks, minimizing network vulnerable time and assisting with failure recovery. Jiarong Xing, Kuo-Feng Hsu, Yiting Xia, Yan Cai 0018, Ying Zhang 0022, Ang Chen 0001 |
EuroSys | 6 |
| 2024 | Rail-only: A Low-Cost High-Performance Network for Training LLMs with Trillion ParametersabstractThis paper presents a low-cost network architecture for training large language models (LLMs) at hyperscale. We study the optimal parallelization strategy of LLMs and propose a novel datacenter network design tailored to LLM's unique communication pattern. We show that LLM training generates sparse communication patterns in the network and, therefore, does not require any-to-any full-bisection network to complete efficiently. As a result, our design eliminates the spine layer in traditional GPU clusters. We name this design a Rail-only network and demonstrate that it achieves the same training performance while reducing the network cost by 38% to 77% and network power consumption by 37% to 75% compared to a conventional GPU datacenter. Our architecture also supports Mixture-of-Expert (MoE) models with all-to-all communication through forwarding, with only 4.1% to 5.6% completion time overhead for all-to-all traffic. We study the failure robustness of Rail-only networks and provide insights into the performance impact of different network and training parameters. Weiyang Wang, Manya Ghobadi, Kayvon Shakeri, Ying Zhang 0022, Naader Hasani |
HOTI | 4 |
| 2024 | Understanding Incast Bursts in Modern DatacentersabstractIn datacenters, common incast traffic patterns are challenging because they violate the basic premise of bandwidth stability on which TCP congestion control convergence is built, overwhelming shallow switch buffers and causing packet losses and high latency. To understand why these challenges remain despite decades of research on datacenter congestion control, we conduct an in-depth investigation into high-degree incasts both in production workloads at Meta and in simulation. In addition to characterizing the bursty nature of these incasts and their impacts on the network, our findings demonstrate the shortcomings of widely deployed window-based congestion control techniques used to address incast problems. Furthermore, we find that hosts associated with a specific application or service exhibit similar and predictable incast traffic properties across hours, pointing the way toward solutions that predict and prevent incast bursts, instead of reacting to them. Christopher Canel, Balasubramanian Madhavan, Srikanth Sundaresan, Neil Spring, Prashanth Kannan, Ying Zhang 0022, Srinivasan Seshan |
IMC | 6 |
| 2024 | Netcastle: Network Infrastructure Testing At Scale
Rob Sherwood, Jinghao Shi, Ying Zhang 0022, Neil Spring, Srikanth Sundaresan, Jasmeet Bagga, Prathyusha Peddi, Vineela Kukkadapu, Rashmi Shrivastava, Manikantan K., Pavan Patil, Srikrishna Gopu, Varun Varadan, Ethan Shi, Hany Morsy, Yuting Bu, Renjie Yang, Rasmus Jonsson, Jesus Jussepen Arredondo, Diana Saha, Sean Choi |
NSDI | 3 |
| 2024 | Transferable Neural WAN TE for Changing TopologiesabstractRecently, researchers have proposed ML-driven traffic engineering (TE) schemes where a neural network model is used to produce TE decisions in lieu of conventional optimization solvers. Unfortunately existing ML-based TE schemes are not explicitly designed to be robust to topology changes that may occur due to WAN evolution, failures or planned maintenance. In this paper, we present HARP, a neural model for TE explicitly capable of handling variations in topology including those not observed in training. HARP is designed with two principles in mind: (i) ensure invariances to natural input transformations (e.g., permutations of node ids, tunnel reordering); and (ii) align neural architecture to the optimization model. Evaluations on a multi-week dataset of a large private WAN show HARP achieves an MLU at most 11% higher than optimal over 98% of the time despite encountering significantly different topologies in testing relative to training data. Further, comparisons with state-of-the-art ML-based TE schemes indicate the importance of the mechanisms introduced by HARP to handle topology variability. Finally, when predicted traffic matrices are provided, HARP outperforms classic optimization solvers achieving a median reduction in MLU of 5 to 10% on the true traffic matrix. Abd AlRhman AlQiam, Yuanjun Yao, Zhaodong Wang, Satyajeet Ahuja, Ying Zhang 0022, Sanjay G. Rao, Bruno Ribeiro 0001, Mohit Tawarmalani |
SIGCOMM | 5 |
| 2024 | NetEdit: An Orchestration Platform for eBPF Network Functions at ScaleabstractManaging the performance of thousands of services across millions of servers demands a networking stack that can dynamically adjust protocol settings to match diverse priorities and network characteristics. Moreover, given the constantly evolving nature of services and their requirements, the set of configurable protocols must remain adaptable. However, current host networking stacks lack the necessary flexibility and adaptability. Although eBPF shows promise in this regard, it lacks essential primitives for efficient development and safe deployment of multiple co-existing services. Theophilus Benson, Prashanth Kannan, Prankur Gupta, Balasubramanian Madhavan, Kumar Saurabh Arora, Martin Lau, Abhishek Dhamija, Rajiv Krishnamurthy, Srikanth Sundaresan, Neil Spring, Ying Zhang 0022 |
SIGCOMM | 12 |
| 2024 | MCCS: A Service-based Approach to Collective Communication for Multi-Tenant CloudabstractPerformance of collective communication is critical for distributed systems. Using libraries to implement collective communication algorithms is not a good fit for a multi-tenant cloud environment because the tenant is not aware of the underlying physical network configuration or how other tenants use the shared cloud network---this lack of information prevents the library from selecting an optimal algorithm. In this paper, we explore a new approach for collective communication that more tightly integrates the implementation with the cloud network instead of the applications. We introduce MCCS, or Managed Collective Communication as a Service, which exposes traditional collective communication abstractions to applications while providing control and flexibility to the cloud provider for their implementations. Realizing MCCS involves overcoming several key challenges to integrate collective communication as part of the cloud network, including memory management of tenant GPU buffers, synchronizing changes to collective communication strategies, and supporting policies that involve cross-layer traffic optimization. Our evaluations show that MCCS improves tenant collective communication performance by up to 2.4× compared to one of the state-of-the-art collective communication libraries (NCCL), while adding more management features including dynamic algorithm adjustment, quality of service, and network-aware traffic engineering. Yechen Xu, Jingrong Chen 0002, Zhaodong Wang, Ying Zhang 0022, Matthew Lentz, Danyang Zhuo |
SIGCOMM | 5 |
| 2023 | Practical Intent-driven Routing Configuration Synthesis
Sivaramakrishnan Ramanathan, Ying Zhang 0022, Mohab Gawish, Yogesh Mundada, Zhaodong Wang, Sangki Yun, Eric Lippert, Walid Taha, Minlan Yu, Jelena Mirkovic |
NSDI | 2 |
| 2023 | TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training Jobs
Weiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi, Dheevatsa Mudigere, Ying Zhang 0022, Anthony Kewitsch |
NSDI | 7 |
| 2023 | EBB: Reliable and Evolvable Express Backbone Network in MetaabstractWe present the design, implementation, evaluation, deployment and production experiences of EBB (Express BackBone), a private WAN (Wide Area Network) connecting Meta's global data centers (DCs). Initiated in 2015, EBB now carries 100% of DC-DC traffic, witnessing remarkable growth over the years. A key design aspect of EBB is its multi-plane architecture, facilitating seamless deployment of a new control plane while ensuring operational simplicity. This architecture allows for efficient failure mitigation, standard maintenance, and capacity expansion by draining one or two planes without impacting service level objectives (SLOs). Another critical design decision is the hybrid model, combining distributed control agents and a central controller. EBB's centralized traffic engineering utilizes an MPLS-TE based solution to allocate paths periodically for different traffic classes based on service requirements, while its distributed control agents enable fast local failure recovery by pre-installing pre-computed backup paths in the data plane. We delve into our eight-year production experience, highlighting the successful deployment of multiple generations of EBB. Marek Denis, Yuanjun Yao, Ashley Hatch, Chiunlin Lim, Shuqiang Zhang, Kyle Sugrue, Henry Kwok, Mikel Jimenez Fernandez, Petr Lapukhov, Sandeep Hebbani, Gaya Nagarajan, Omar Baldonado, Lixin Gao 0001, Ying Zhang 0022 |
SIGCOMM | 15 |
| 2023 | FlexWAN: Software Hardware Co-design for Cost-Effective and Resilient Optical BackbonesabstractThe rising demand for WAN capacity driven by the rapid growth of inter-data center traffic poses new challenges for costly optical networks. Today cloud providers rely on fixed optical backbones, where all hardware devices operate on a rigid spectrum grid, leading to the waste of expensive optical resources and subpar performance in handling failures. In this paper, we introduce FlexWAN, a novel flexible WAN infrastructure designed to provision cost-effective WAN capacity while ensuring resilience to optical failures. FlexWAN achieves this by incorporating spacing-variable hardware at the optical layer, enabling the generated wavelength to optimize the utilization of limited spectrum resources for the WAN capacity. The configuration of spacing-variable hardware in a multi-vendor optical backbone presents challenges related to spectrum management. To address this, FlexWAN leverages a centralized controller to achieve coordinated control of network-wide optical devices in a vendor-agnostic manner. Moreover, the flexibility at the optical layer introduces new algorithmic problems. FlexWAN formulates the problem of provisioning WAN capacity with the goal of minimizing hardware costs. We evaluate the system performance in production and share insights from years of production experience. Compared to existing optical backbones, FlexWAN can save at least 57% of transponders and reduce 36% of spectrum usage while continuing to meet up to 8× the present-day demands using existing hardware and fiber deployments. FlexWAN further incorporates failure resilience that revives 15% more bandwidth capacity in the overloaded optical backbone. Congcong Miao, Zhizhen Zhong, Ying Zhang 0022, Kunling He, Fangchao Li, Minggang Chen, Xiang Li 0223, Zekun He, Xianneng Zou, Jilong Wang 0001 |
SIGCOMM | 3 |
| 2023 | Klotski: Efficient and Safe Network Migration of Large Production DatacentersabstractThis paper presents the design, implementation, evaluation, and deployment of Meta's production network migration system. We first introduce the network migration problem for large-scale production datacenter networks (DCNs). A network migration task at Meta touches as many as hundreds of switches and tens of thousands of circuits per datacenter (DC), and involves physical deployment work on site that can last months. We describe real-world migration challenges, covering complex and evolving DCN architectures and operational constraints. We mathematically formalize the problem of generating efficient and safe migration plans, and exploit the inherent symmetry and locality of DCN topologies to prune the search space. We design an ordering-agnostic compact topology representation to eliminate redundant satisfiability checking, and apply the A* algorithm with a domain-specific priority function to find the optimal plan. Evaluation results on a range of production migration cases show that Klotski reduces the time to find optimal migration plans by up to 381× compared to prior solutions. We hope by introducing the problem and sharing our deployment experience, this work can provide a useful context for network migration in the real world and inspire future research. Xiaoxiang Zhang, Ying Zhang 0022, Zhaodong Wang, Yuandong Tian, Alex Nikulkov, Joao Ferreira, Xuanzhe Liu, Xin Jin 0008 |
SIGCOMM | 4 |
| 2022 | NQ/ATP: Architectural Support for Massive Aggregate Queries in Data Center NetworksabstractNetwork queries become increasingly challenging for online service providers with massive network devices and massive network queries due to the tradeoff between system scale and query granularity. We re-architect the traditional three-tier architecture, i.e., data collection, data storage, and data query, for aggregate queries, and build a system named NQ/ATP. NQ/ATP offloads the aggregation operation in network queries onto network switches, which accelerates the query execution and frees up network resources. NQ/ATP further devises a route learning mechanism, query hierarchy load balancing policy, and hierarchy clustering mechanism to save forwarding table entries on switches, which better supports massive queries. The evaluation shows that NQ/ATP can support network aggregate queries with higher capacity, less traffic volume, finer granularity, and better scalability than traditional three-tier polling architectures. The three optimizations can effectively reduce the forwarding table usage by up to 97.55%. Wenfei Wu, Shan-Hsiang Shen, Ying Zhang 0022 |
IWQoS | 4 |
| 2022 | Evolvable Network Telemetry at Facebook
Yang Zhou 0008, Ying Zhang 0022, Minlan Yu, Dexter Cao, Yu-Wei Eric Sung, Starsky H. Y. Wong |
NSDI | 2 |
| 2022 | Network entitlement: contract-based network sharing with agility and SLO guaranteesabstractThis paper presents Meta's Production Wide Area Network (WAN) Entitlement solution used by thousands of Meta's services to share the network safely and efficiently. We first introduce the Network Entitlement problem, i.e., how to share WAN bandwidth across services with flexibility and SLO guarantees. We present a new abstraction entitlement contract, which is stable, simple, and operationally friendly. The contract defines services' network quota and is set up between the network team and services teams to govern their obligations. Our framework includes two key parts: (1) an entitlement granting system that establishes an agile contract while achieving network efficiency and meeting long-term SLO guarantees, and (2) a large-scale distributed run-time enforcement system that enforces the contract on the production traffic. We demonstrate its effectiveness through extensive simulations and real-world end-to-end tests. The system has been deployed and operated for over two years in production. We hope that our years of experience provide a new angle to viewing WAN network sharing in production and will inspire follow-up research. Satyajeet Ahuja, Vinayak Dangui, Kirtesh Patil, Manikandan Somasundaram, Mario A. Sánchez, Guanqing Yan, Mohammad Noormohammadpour, Alaleh Razmjoo, Grace Smith, Abhinav Triguna, Soshant Bali, Yuxiang Xiang, Prabhakaran Ganesan, Mikel Jimenez Fernandez, Petr Lapukhov, Guyue Liu, Ying Zhang 0022 |
SIGCOMM | 20 |
| 2022 | Flash: fast, consistent data plane verification for large-scale network settingsabstractData plane verification can be an important technique to reduce network disruptions, and researchers have recently made significant progress in achieving fast data plane verification. However, as we apply existing data plane verification techniques to large-scale networks, two problems appear due to extremes. First, existing techniques cannot handle too-fast arrivals, which we call update storms, when a large number of data plane updates must be processed in a short time. Second, existing techniques cannot handle well too-slow arrivals, which we call long-tail update arrivals, when the updates from a number of switches take a long time to arrive. Shenshen Chen, Kai Gao 0001, Qiao Xiang, Ying Zhang 0022, Yang Richard Yang |
SIGCOMM | 5 |
| 2021 | NFD: Using Behavior Models to Develop Cross-Platform Network FunctionsabstractNFV ecosystem is flourishing and more and more NF platforms appear, but this makes NF vendors difficult to deliver NFs rapidly to diverse platforms. We propose an NF development framework named NFD for cross-platform NF development. NFD's main idea is to decouple the functional logic from the platform logic -it provides a platform-independent language to program NFs' behavior models, and a compiler with interfaces to develop platform-specific plugins. By enabling a plugin on the compiler, various NF models would be compiled to executables integrated with the target platform. We prototype NFD, build 14 NFs, and support 6 platforms (standard Linux, OpenNetVM, GPU, SGX, DPDK, OpenNF). Our evaluation shows that NFD can save development workload for cross-platform NFs and output valid and performant NFs. Hongyi Huang, Wenfei Wu, Yongchao He, Bangwen Deng, Ying Zhang 0022, Yongqiang Xiong, Guo Chen 0001, Yong Cui 0001, Peng Cheng 0005 |
INFOCOM | 5 |
| 2021 | Flow Algebra: Towards an Efficient, Unifying Framework for Network Management TasksabstractA modern network needs to conduct a diverse set of tasks, and the existing approaches focus on developing specific tools for specific tasks, resulting in increasing complexity and lacking reusability. In this paper, we propose Flow Algebra as a unifying, easy-to-use framework to accomplish a large set of network management tasks. Based on the observation that relational databases based on relational algebra are well understood and widely used as a unifying framework for data management, we develop flow algebra based on relational algebra. On the other hand, flow tables, which are the fundamental data specifying the state of a network, cannot be stored in traditional relations, because of fundamental features such as wildcard and priorities. We define flow algebra based on novel, generalized relational operations that use equivalency to achieve efficient, unifying data store, query, and manipulation of both flow tables and traditional relations. We realize flow algebra with FlowDB and demonstrate its ease of use on diverse tasks. We further demonstrate that generality and ease-of-use do not need to come with a performance penalty. For example, for the well-studied network verification task, our system outperforms two state-of-the-art network verification engines, NoD and HSA, in their targeted domain, by 55x. Christopher Leet, Robert Soulé, Yang Richard Yang, Ying Zhang 0022 |
INFOCOM | 4 |
| 2021 | A Social Network Under Social Distancing: Risk-Driven Backbone Management During COVID-19 and Beyond
Yiting Xia, Ying Zhang 0022, Zhizhen Zhong, Guanqing Yan, Chiunlin Lim, Satyajeet Ahuja, Soshant Bali, Alexander Nikolaidis, Kimia Ghobadi, Manya Ghobadi |
NSDI | 2 |
| 2021 | Capacity-efficient and uncertainty-resilient backbone network planning with hoseabstractThis paper presents Facebook's design and operational experience of a Hose-based backbone network planning system. This initial adoption of the Hose model in network planning is driven by the capacity and demand uncertainty pressure of backbone expansion. Since the Hose model abstracts the aggregated traffic demand per site, peak traffic flows at different times can be multiplexed to save capacity and buffer traffic spikes. Our core design involves heuristic algorithms to select Hose-compliant traffic matrices and cross-layer optimization between the optical and IP networks. We evaluate the system performance in production and share insights from years of production experience. Hose-based network planning can save 17.4% capacity and drops 75% less traffic under fiber cuts. As the first study of Hose in network planning, our work has the potential to inspire follow-up research. Satyajeet Ahuja, Vinayak Dangui, Soshant Bali, Abishek Gopalan, Petr Lapukhov, Yiting Xia, Ying Zhang 0022 |
SIGCOMM | 9 |
| 2021 | ARROW: restoration-aware traffic engineeringabstractFiber cut events reduce the capacity of wide-area networks (WANs) by several Tbps. In this paper, we revive the lost capacity by reconfiguring the wavelengths from cut fibers into healthy fibers. We highlight two challenges that made prior solutions impractical and propose a system called Arrow to address them. First, our measurements show that contrary to common belief, in most cases, the lost capacity is only partially restorable. This poses a cross-layer challenge from the Traffic Engineering (TE) perspective that has not been considered before: “Which IP links should be restored and by how much to best match the TE objective?” To address this challenge, Arrow's restoration-aware TE system takes a set of partial restoration candidates (that we call LotteryTickets) as input and proactively finds the best restoration plan. Second, prior work has not considered the reconfiguration latency of amplifiers. However, in practical settings, amplifiers add tens of minutes of reconfiguration delay. To enable fast and practical restoration, Arrow leverages optical noise loading and bypasses amplifier reconfiguration altogether. We evaluate Arrow using large-scale simulations and a testbed. Our testbed demonstrates Arrow's end-to-end restoration latency is eight seconds. Our large-scale simulations compare Arrow to the state-of-the-art TE schemes and show it can support 2.0x--2.4x more demand without compromising 99.99% availability. Zhizhen Zhong, Manya Ghobadi, Alaa Khaddaj, Jonathan Leach, Yiting Xia, Ying Zhang 0022 |
SIGCOMM | 6 |
| 2021 | Network planning with deep reinforcement learningabstractNetwork planning is critical to the performance, reliability and cost of web services. This problem is typically formulated as an Integer Linear Programming (ILP) problem. Today's practice relies on hand-tuned heuristics from human experts to address the scalability challenge of ILP solvers. Satyajeet Ahuja, Yuandong Tian, Ying Zhang 0022, Xin Jin 0008 |
SIGCOMM | 5 |
| 2021 | Providing Bandwidth Guarantees, Work Conservation and Low Latency Simultaneously in the CloudabstractToday's cloud is shared among multiple tenants running different applications, and a desirable multi-tenant datacenter network infrastructure should provide bandwidth guarantees for throughput-intensive applications, low latency for latency-sensitive short messages, as well as work conservation to fully utilize the network bandwidth. Despite significant efforts in recent years, none of them can achieve these three properties simultaneously. In this paper, we identify the key deficiency of prior solutions and use this insight to motivate our design of Trinity-a simple, practical yet effective solution that achieves bandwidth guarantees, work conservation and low latency simultaneously in the cloud. We implement Trinity using existing commodity hardwares and demonstrate its superior performance over prior solutions using testbed experiments. Shuihai Hu, Wei Bai 0001, Kai Chen 0005, Chen Tian 0001, Ying Zhang 0022 |
IEEE Trans. Cloud Comput. | 5 |
| 2020 | Concury: a fast and light-weight software cloud load balancerabstractA load balancer (LB) is a vital network function for cloud services to balance the load amongst resources. Stateful software LBs that run on commodity servers provide flexibility, cost-efficiency, and packet consistency. However, current designs have two main limitations: 1) states are stored as digests, which may cause packet inconsistency due to digest collisions; 2) the data plane needs to update for every new connection, and frequent updates hurt throughput and packet consistency. In this work, we present a new software stateful LB called Concury, which is the first solution to solve these problems. The key innovation of Concury is a new method to maintain large network states with frequent connection arrivals, which is succinct in memory cost, consistent under network changes, and incurs low update cost. The evaluation results show that the Concury algorithm provides 4x throughput and consumes less memory compared to other LB algorithms, while providing weighted load balancing and false-hit freedom, for both real and synthetic data center traffic. We implement Concury and evaluate it in two real networks. It achieves 67.2 Gbps single-thread throughput on a cheap desktop computer in 100GbE. Shouqian Shi, Ye Yu 0001, Minghao Xie, Xin Li 0057, Ying Zhang 0022, Chen Qian 0001 |
SoCC | 6 |
| 2019 | Alembic: Automated Model Inference for Stateful Network Functions
Soo-Jin Moon, Jeffrey Helt, Yves Bieri, Sujata Banerjee, Vyas Sekar, Wenfei Wu, Mihalis Yannakakis, Ying Zhang 0022 |
NSDI | 9 |
| 2019 | Special Issue on Artificial Intelligence and Machine Learning for Networking and CommunicationsabstractResearch in large-scale networking systems has been shaped and will continue to be guided by specific characteristics of applications and the underlying platforms and infrastructures. On the one hand, applications are growing at an accelerated pace, which is fundamentally unpredictable in both breadth and depth. On the other hand, the underlying networking has been the focus of a huge transformation enabled by new models resulting from virtualization and cloud computing. This has led to a number of novel architectures supported by emerging technologies such as Software-Defined Networking (SDN), Network Function Virtualization (NFV), and more recently, edge cloud and fog networking, or network slicing[1],[2]. This evolution towards enhanced design opportunities along with increasing complexity in networking and its applications has fueled the need for improved network automation in agile infrastructures. At the same time, their complexity has dramatically increased. The networking dynamics have had the effect of making it even more important and challenging to design scalable network measurement and analysis techniques and associated tools. Critical applications such as resource allocation, network monitoring, security enforcement, or dynamic network management require real-time mechanisms for online analysis as well as efficient techniques for offline deep analysis of massive historical data. Prosper Chemouil, Pan Hui 0001, Wolfgang Kellerer, Yong Li 0008, Rolf Stadler, Dacheng Tao, Yonggang Wen 0001, Ying Zhang 0022 |
IEEE J. Sel. Areas Commun. | 8 |
| 2019 | Virtual Machine Migration Planning in Software-Defined NetworksabstractLive migration is a key technique for virtual machine (VM) management in data center networks, which enables flexibility in resource optimization, fault tolerance, and load balancing. Despite its usefulness, the live migration still introduces performance degradations during the migration process. Thus, there has been continuous efforts in reducing the migration time in order to minimize the impact. From the network's perspective, the migration time is determined by the amount of data to be migrated and the available bandwidth used for such transfer. In this paper, we examine the problem of how to schedule the migrations and how to allocate network resources for migration when multiple VMs need to be migrated at the same time. We consider the problem in the Software-defined Network (SDN) context since it provides flexible control on routing. More specifically, we propose a method that computes the optimal migration sequence and network bandwidth used for each migration. We formulate this problem as a mixed integer programming, which is NP-hard. To make it computationally feasible for large scale data centers, we propose an approximation scheme via linear approximation plus fully polynomial time approximation, and obtain its theoretical performance bound and computational complexity. Through extensive simulations, we demonstrate that our fully polynomial time approximation (FPTA) algorithm has a good performance compared with the optimal solution of the primary programming problem and two state-of-the-art algorithms. That is, our proposed FPTA algorithm approaches to the optimal solution of the primary programming problem with less than 10 percent variation and much less computation time. Meanwhile, it reduces the total migration time and service downtime by up to 40 and 20 percent compared with the state-of-the-art algorithms, respectively. Huandong Wang, Yong Li 0008, Ying Zhang 0022, Depeng Jin |
IEEE Trans. Cloud Comput. | 3 |
| 2018 | SENSS Against Volumetric DDoS AttacksabstractVolumetric distributed denial-of-service (DDoS) attacks can bring any network to a halt. Because of their distributed nature and high volume, the victim often cannot handle these attacks alone and needs help from upstream ISPs. Today's Internet has no automated mechanism for victims to ask ISPs for help in attack handling and ISPs themselves do not offer such services. We propose SENSS, a security service for collaborative mitigation of volumetric DDoS attacks. SENSS enables the victim of an attack to request attack monitoring and filtering on demand, and to pay for the services rendered. Requests can be sent both to the immediate and to remote ISPs, in an automated and secure manner, and can be authenticated by these ISPs, without having prior trust with the victim. Simple and generic SENSS APIs enable victims to build custom detection and mitigation approaches against a variety of DDoS attacks. SENSS is deployable with today's infrastructure, and it has strong economic incentives both for ISPs and for the attack victims. It is also very effective in sparse deployment, offering full protection to direct customers of early adopters, and considerable protection to remote victims when deployed strategically. Deployment on the largest 1% of ISPs protects not just direct customers of these ISPs, but everyone on the Internet, from 90% of volumetric DDoS attacks. Sivaramakrishnan Ramanathan, Jelena Mirkovic, Minlan Yu, Ying Zhang 0022 |
ACSAC | 4 |
| 2018 | NetCP: Consistent, Non-Interruptive and Efficient Checkpointing and Rollback of SDNabstractNetwork failures are inevitable due to its increasing complexity, which significantly hampers system availability and performance. While adopting checkpointing and rollback recovery protocols (C/R for abbreviation) from distributed systems into computer networks is promising, several specific challenges appear as we design a C/R system for Software-Defined Networks (SDN). The C/R should be coordinated with other applications in the SDN controller, each individual switch C/R should not interrupt traffic traversing it, and SDN controller C/R faces the challenge of time and space overhead. We propose a C/R framework for SDN, named NetCP. NetCP coordinates C/R and other applications to get consistent global checkpoints, it leverages redundant forwarding tables in SDN switches for C/R so as to avoid interrupting traversing traffic, and it analyzes the dependencies between controller applications to make minimal C/R decision. We have implemented NetCP in a prototype system using the current standard SDN tools and demonstrate that it achieves consistency, non-interruption, and efficiency with negligible overhead. Ye Yu 0001, Chen Qian 0001, Wenfei Wu, Ying Zhang 0022 |
IWQoS | 4 |
| 2018 | FBOSS: building switch software at scaleabstractThe conventional software running on network devices, such as switches and routers, is typically vendor-supplied, proprietary and closed-source; as a result, it tends to contain extraneous features that a single operator will not most likely fully utilize. Furthermore, cloud-scale data center networks often times have software and operational requirements that may not be well addressed by the switch vendors. Sean Choi, Boris Burkov, Alex Eckert, Tian Fang, Saman Kazemkhani, Rob Sherwood, Ying Zhang 0022, Hongyi Zeng |
SIGCOMM | 7 |
| 2018 | Dragon: Scalable, Flexible, and Efficient Traffic Engineering in Software Defined ISP NetworksabstractTo optimize network cost, performance, and reliability, SDN advocates for centralized traffic engineering (TE) that enables more efficient path and egress point selection as well as bandwidth allocation. In this paper, we argue SDN-based TE for ISP networks can be very challenging. First, ISP networks often are very large in size, imposing significant scalability challenges to the centralized TE. Second, ISP networks usually have diverse types of links, switches, and cost models, leading to a complex combination of optimizations. Third, ISP networks not only have many choices of internal paths but also include rich selections of egress points and interdomain routes, unlike cloud/enterprise networks. To overcome these challenges, we present a novel TE application framework, called Dragon, for existing SDN control planes. To address the scalability challenge, Dragon consists of hierarchical and recursive TE algorithms and mechanisms that divide flow optimization problems into subtasks and execute them in parallel. Further, Dragon allows ISPs to express diverse objectives for different parts of their network. Finally, we extend Dragon to jointly optimize the selection of intradomain and interdomain paths. Using extensive evaluation on real topologies and prototyping with SDN controller and switches, we demonstrate that Dragon outperforms existing TE methods both in speed and optimality. Mehrdad Moradi, Ying Zhang 0022, Z. Morley Mao, Ravi Manghirmalani |
IEEE J. Sel. Areas Commun. | 2 |
| 2017 | Low Latency Software Rate Limiters for Cloud NetworksabstractA lot of recent work has focused on reducing in network queueing latency in datacenter networks. In this paper, we focus on a less explored topic --- latency increases caused by queueing in rate limiters on the end-host. First, we show that latency can be increased by an order of magnitude by rate limiters in cloud networks. To solve this problem, we extend ECN marking into rate limiters and use a datacenter congestion control algorithm --- DCTCP. Unfortunately, while this reduces latency, it also leads to throughput oscillation. Thus, this solution is not sufficient. In this paper, we also analyze the specific reasons that ECN marking in software rate limiters leads to the throughput oscillation problem. Finally, we propose two potential solutions to design software rate limiters that can achieve stable high throughput and low latency. Keqiang He, Weite Qin, Wenfei Wu, Tian Pan 0001, Chengchen Hu, Jiao Zhang 0002, Brent E. Stephens, Aditya Akella, Ying Zhang 0022 |
APNet | 11 |
| 2017 | Supporting Diverse Dynamic Intent-based Policies using JanusabstractExisting network policy abstractions handle basic group based reachability and access control list based security policies. However, QoS policies as well as dynamic policies are also important and not representing them in the high level policy abstraction poses serious limitations. At the same time, efficiently configuring and composing group based QoS and dynamic policies present significant technical challenges, such as (a) maintaining group granularity during configuration, (b) dealing with network-bandwidth contention among policies from distinct writers and (c) dealing with multiple path changes corresponding to dynamically changing policies, group membership and end-point mobility. In this paper we propose Janus, a system which makes two major contributions. First, we extend the prior policy graph abstraction model to represent complex QoS and dynamic tateful/temporal policies. Second, we convert the policy configuration problem into an optimization problem with the goal of maximizing the number of satisfied and configured policies, and minimizing the number of path changes under dynamic environments. To solve this, Janus presents several novel heuristic algorithms. We evaluate our system using a diverse set of bandwidth policies and network topologies. Our experiments demonstrate that Janus can achieve near-optimal solutions in a reasonable amount of time. Anubhavnidhi Abhashkumar, Joon-Myung Kang, Sujata Banerjee, Aditya Akella, Ying Zhang 0022, Wenfei Wu |
CoNEXT | 5 |
| 2017 | An empirical characterization of IFTTT: ecosystem, usage, and performanceabstractIFTTT is a popular trigger-action programming platform whose applets can automate more than 400 services of IoT devices and web applications. We conduct an empirical study of IFTTT using a combined approach of analyzing data collected for 6 months and performing controlled experiments using a custom testbed. We profile the interactions among different entities, measure how applets are used by end users, and test the performance of applet execution. Overall we observe the fast growth of the IFTTT ecosystem and its increasing usage for automating IoT-related tasks, which correspond to 52% of all services and 16% of the applet usage. We also observe several performance inefficiencies and identify their causes. Xianghang Mi, Feng Qian 0001, Ying Zhang 0022, XiaoFeng Wang 0001 |
Internet Measurement Conference | 3 |
| 2017 | SLA-verifier: Stateful and quantitative verification for service chainingabstractNetwork verification has been recently proposed to detect network misconfigurations. Existing work focuses on the reachability. This paper proposes a framework that verifies the Service Level Agreement (SLA) compliance of the network using static verification. This work proposes a quantitative model and a set of algorithms for verifying performance properties of a network with switches and middleboxes, i.e., service chains. We develop SLA-Verifier and evaluate its efficiency using simulation on real-world data and testbed experiments. To improve the SLA violation detection accuracy, our system uses verification results to optimize online monitoring. Ying Zhang 0022, Wenfei Wu, Sujata Banerjee, Joon-Myung Kang, Mario A. Sánchez |
INFOCOM | 1 |
| 2017 | Improve Service Chaining Performance with Optimized Middlebox PlacementabstractPrevious works have proposed various approaches to implement service chaining by routing traffic through the desired middleboxes according to pre-defined policies. However, no matter what routing scheme is used, the performance of service chaining depends on where these middleboxes are placed. Thus, in this paper, we study middlebox placement problem, i.e., given network information and policy specifications, we attempt to determine the optimal locations to place the middleboxes so that the performance is optimized. The performance metrics studied in this paper include the end-to-end delay and the bandwidth consumption, which cover both users’ and network providers’ interests. We first formulate it as 0-1 programming problem, and prove it is NP-hard. We then propose two heuristic algorithms to obtain the sub-optimal solutions. The first algorithm is a greedy algorithm, and the second algorithm is based on simulated annealing. Through extensive simulations, we show that in comparison with a baseline algorithm, the proposed algorithms can reduce 22 percent end-to-end delay and save 38 percent bandwidth consumption on average. The formulation and proposed algorithms have no special assumption on network topology or policy specifications, therefore, they have broad range of applications in various types of networks such as enterprise, data center and broadband access networks. Jiaqiang Liu, Yong Li 0008, Ying Zhang 0022, Li Su 0001, Depeng Jin |
IEEE Trans. Serv. Comput. | 3 |
| 2016 | Automatic Synthesis of NF Models by Program AnalysisabstractNetwork functions (NFs), like firewall, NAT, IDS, have been widely deployed in today’s modern networks. However, currently there is no standard specification or modeling language that can accurately describe the complexity and diversity of different NFs. Recently there have been research efforts to propose NF models. However, they are often generated manually and thus error-prone. This paper proposes a method to automatically synthesize NF models via program analysis. We develop a tool called NFactor, which conducts code refactoring and program slicing on NF source code, in order to generate its forwarding model. We demonstrate its usefulness on two NFs and evaluate its correctness. A few applications of NFactor are described, including network verification. Wenfei Wu, Ying Zhang 0022, Sujata Banerjee |
HotNets | 2 |
| 2016 | Providing bandwidth guarantees, work conservation and low latency simultaneously in the cloudabstractToday's cloud is shared among multiple tenants running different applications, and a desirable multi-tenant datacenter network infrastructure should provide bandwidth guarantees for throughput-intensive applications, low latency for latency-sensitive short messages, as well as work conservation to fully utilize the network bandwidth. Despite significant efforts in recent years, none of them can achieve these three properties simultaneously. In this paper, we identify the key deficiency of prior solutions and use this insight to motivate our design of Trinity - a simple, practical yet effective solution that achieves bandwidth guarantees, work conservation and low latency simultaneously in the cloud. We implement Trinity using existing commodity hardwares and demonstrate its superior performance over prior solutions using testbed experiments. Shuihai Hu, Wei Bai 0001, Kai Chen 0005, Chen Tian 0001, Ying Zhang 0022 |
INFOCOM | 5 |
| 2015 | Virtual machine migration planning in software-defined networksabstractLive migration is a key technique for virtual machine (VM) management in data center networks, which enables flexibility in resource optimization, fault tolerance, and load balancing. Despite its usefulness, the live migration still introduces performance degradations during the migration process. Thus, there has been continuous efforts in reducing the migration time in order to minimize the impact. From the network's perspective, the migration time is determined by the amount of data to be migrated and the available bandwidth used for such transfer. In this paper, we examine the problem of how to schedule the migrations and how to allocate network resources for migration when multiple VMs need to be migrated at the same time. We consider the problem in the Software-defined Network (SDN) context since it provides flexible control on routing. More specifically, we propose a method that computes the optimal migration sequence and network bandwidth used for each migration. We formulate this problem as a mixed integer programming, which is NP-hard. To make it computationally feasible for large scale data centers, we propose an approximation scheme via linear approximation plus fully polynomial time approximation, and obtain its theoretical performance bound. Through extensive simulations, we demonstrate that our fully polynomial time approximation (FPTA) algorithm has a good performance compared with the optimal solution and two state-of-the-art algorithms. That is, our proposed FPTA algorithm approaches to the optimal solution with less than 10% variation and much less computation time. Meanwhile, it reduces the total migration time and the service downtime by up to 40% and 20% compared with the state-of-the-art algorithms, respectively. Huandong Wang, Yong Li 0008, Ying Zhang 0022, Depeng Jin |
INFOCOM | 3 |
| 2015 | Network Policy Whiteboarding and CompositionabstractWe present Policy Graph Abstraction (PGA) that graphically expresses network policies and service chain requirements, just as simple as drawing whiteboard diagrams. Different users independently draw policy graphs that can constrain each other. PGA graph clearly captures user intents and invariants and thus facilitates automatic composition of overlapping policies into a coherent policy. Jeongkeun Lee, Joon-Myung Kang, Chaithan Prakash, Yoshio Turner, Aditya Akella, Charles Clark, Yadi Ma, Puneet Sharma 0001, Ying Zhang 0022 |
SIGCOMM | 9 |
| 2015 | PGA: Using Graphs to Express and Automatically Reconcile Network PoliciesabstractSoftware Defined Networking (SDN) and cloud automation enable a large number of diverse parties (network operators, application admins, tenants/end-users) and control programs (SDN Apps, network services) to generate network policies independently and dynamically. Yet existing policy abstractions and frameworks do not support natural expression and automatic composition of high-level policies from diverse sources. We tackle the open problem of automatic, correct and fast composition of multiple independently specified network policies. We first develop a high-level Policy Graph Abstraction (PGA) that allows network policies to be expressed simply and independently, and leverage the graph structure to detect and resolve policy conflicts efficiently. Besides supporting ACL policies, PGA also models and composes service chaining policies, i.e., the sequence of middleboxes to be traversed, by merging multiple service chain requirements into conflict-free composed chains. Our system validation using a large enterprise network policy dataset demonstrates practical composition times even for very large inputs, with only sub-millisecond runtime latencies. Chaithan Prakash, Jeongkeun Lee, Yoshio Turner, Joon-Myung Kang, Aditya Akella, Sujata Banerjee, Charles Clark, Yadi Ma, Puneet Sharma 0001, Ying Zhang 0022 |
SIGCOMM | 10 |
| 2014 | User mobility from the view of cellular data networksabstractUnderstanding the user mobility is essential to resource optimization and algorithm evaluation in mobile networks, such as network planning, content distribution, and evaluation of hand-over mechanisms. Existing human mobility models focus on extracting mobility patterns from Call Detail Records (CDRs) or WiFi traces. While the former only captures movements during phone calls, the latter does not provide direct answers to mobility of cellular network users in a large scale. In this paper, we take the first step to investigate if the mobility properties derived from cellular data traffic is different from the previous findings using other data source, especially the commonly used CDR based approach. We present a comprehensive characterization of the mobility patterns from the cellular data networks' perspective, using a set of systematic methods. We find that the data network records can provide finer granularity of location and movement information. Three different temporal movement patterns are identified. Furthermore, we propose a new method for predicting future application usage given the mobility patterns and show promising results. Ying Zhang 0022 |
INFOCOM | 1 |
| 2014 | Understanding HTTP flow rates in cellular networksabstractData traffic in cellular networks increased tremendously over the past few years and this growth is predicted to continue over the next few years. Due to differences in access technology and user behavior, the characteristics of cellular traffic can differ from existing results for wireline traffic. In this study we focus on understanding the flow rates and on the relationship between the rates and other flow properties by analyzing packet level traces collected in a large cellular network. To understand the limiting factors of the flow rates, we further analyze the underlying causes behind the observed rates, e.g., network congestion, access link or end host configuration. Our study extends other related work by conducting the analysis from a unique dimension, the comparison with traffic in wired networks, to reveal the unique properties of cellular traffic. We find that they differ in variability and in the dominant rate limiting factors. Ying Zhang 0022, Åke Arvidsson, Matti Siekkinen, Guillaume Urvoy-Keller |
Networking | 1 |
| 2014 | SENSS: observe and control your own traffic in the internetabstractWe propose a new software-defined security service -- SENSS -- that enables a victim network to request services from remote ISPs for traffic that carries source IPs or destination IPs from this network's address space. These services range from statistics gathering, to filtering or quality of service guarantees, to route reports or modifications. The SENSS service has very simple, yet powerful, interfaces. This enables it to handle a variety of data plane and control plane attacks, while being easily implementable in today's ISP. Through extensive evaluations on realistic traffic traces and Internet topology, we show how SENSS can be used to quickly, safely and effectively mitigate a variety of large-scale attacks that are largely unhandled today. Abdulla Alwabel, Minlan Yu, Ying Zhang 0022, Jelena Mirkovic |
SIGCOMM | 3 |
| 2013 | An adaptive flow counting method for anomaly detection in SDNabstractThe accuracy and granularity of network flow measurement play a critical role in many network management tasks, especially for anomaly detection. Despite its important, traffic monitoring often introduces overhead to the network, thus, operators have to employ sampling and aggregation to avoid overloading the infrastructure. However, such sampled and aggregated information may affect the accuracy of traffic anomaly detection. In this work, we propose a novel method that performs adaptive zooming in the aggregation of flows to be measured. In order to better balance the monitoring overhead and the anomaly detection accuracy, we propose a prediction based algorithm that dynamically change the granularity of measurement along both the spatial and the temporal dimensions. To control the load on each individual switch, we carefully delegate monitoring rules in the network wide. Using real-world data and three simple anomaly detectors, we show that the adaptive based counting can detect anomalies more accurately with less overhead. Ying Zhang 0022 |
CoNEXT | 1 |
| 2013 | StEERING: A software-defined networking for inline service chainingabstractNetwork operators are faced with the challenge of deploying and managing middleboxes (also called inline services) such as firewalls within their broadband access, datacenter or enterprise networks. Due to the lack of available protocols to route traffic through middleboxes, operators still rely on error-prone and complex low-level configurations to coerce traffic through the desired set of middleboxes. Built upon the recent software-defined networking (SDN) architecture and OpenFlow protocol, this paper proposes StEERING, short for SDN inlinE sERvices and forwardlNG. It is a scalable framework for dynamically routing traffic through any sequence of middleboxes. With simple centralized configuration, StEERING can explicitly steer different types of flows through the desired set of middleboxes, scaling at the level of per-subscriber and per-application policies. With its capability to support flexible routing, we further propose an algorithm to select the best locations for placing services, such that the performance is optimized. Overall, StEERING allows network operators to monetize their middlebox deployment in new ways by allowing subscribers flexibly to select available network services. Ying Zhang 0022, Neda Beheshti, Ludovic Béliveau, Geoffrey Lefebvre, Ravi Manghirmalani, Ramesh Mishra, Ritun Patney, Meral Shirazipour, Ramesh Subrahmaniam, Catherine Truchan, Mallik Tatipamula |
ICNP | 1 |
| 2013 | Detecting user dissatisfaction and understanding the underlying reasonsabstractQuantifying quality of experience for network applications is challenging as it is a subjective metric with multiple dimensions such as user expectation, satisfaction, and overall experience. Today, despite various techniques to support differentiated Quality of Service (QoS), the operators still lack of automated methods to translate QoS to QoE, especially for general web applications. Åke Arvidsson, Ying Zhang 0022 |
SIGMETRICS | 2 |
| 2012 | Fast failover for control traffic in Software-defined NetworksabstractThe Software-defined Network (SDN) design decouples forwarding and control planes, and runs the controlling functions on servers that might be in different physical locations from the forwarding elements. Such separation introduces new challenges to the network resiliency, because disconnection between switches and the controller could disable the forwarding plane. In this work, we analyze resiliency of the connection between control and forwarding planes in SDN. We propose algorithms to improve this resiliency by maximizing the possibility of fast failover-which we achieve through resilience-aware controller placement and control-traffic routing in the network. Neda Beheshti, Ying Zhang 0022 |
GLOBECOM | 2 |
| 2012 | Studying Impacts of Prefix Interception Attack by Exploring BGP AS-PATH PrependingabstractThe AS path prep ending approach in BGP is commonly used to perform inter-domain traffic engineering, such as inbound traffic load-balancing for multi-homed ASes. It artificially increases the length of the AS level path in BGP announcements by inserting its local AS number multiple times into outgoing announcements. In this work, we study how the AS path prep ending mechanism can be exploited to launch a BGP prefix interception attack. Our work is motivated by a recent routing anomaly related to AS Path prepending behavior, i.e., Facebook's traffic being redirected to Korea and China due to a shorter path with fewer prep ending ASNs. In order to measure the possible impact of the attack, we develop a simulator to quantify the damage of the attack under a diverse set of attacker/victim combinations. Our main contribution is to quantify how many ASes may be susceptible to the attack, and analyze how effective the attack may be through simulation. Furthermore, we propose an algorithm to detect the interception attack by exploiting inconsistencies via collaborative monitoring from multiple vantage points. Our evaluation shows up to 99% accuracy with 150 vantage points. Ying Zhang 0022, Makan Pourzandi |
ICDCS | 1 |
| 2012 | The Freshman Handbook: A Hint for Server Placement in Online Social Network ServicesabstractFor new social service providers (freshmen), it is critical to determine where to deploy the computational resources to best accommodate future client requests. Existing proposals on server placement rely on collecting and analyzing request history from servers that are already running, which are not so useful to those starting new online social network services (OSNs). In this work, we aim at helping new OSN providers with intelligent server placement by exploring available public information from existing social network communities. We explore the commonality between the selected set of server locations from multiple OSNs and utilize such similarity for future OSNs to select their server locations. The similarity is ultimately due to the fact that the underlying human relationship is relatively consistent across OSNs. Ying Zhang 0022, Du Li, Mallik Tatipamula |
ICPADS | 1 |
| 2012 | Formal Verification of Security Preservation for Migrating Virtual Machines in the Cloud
Yosr Jarraya, Arash Eghtesadi, Mourad Debbabi, Ying Zhang 0022, Makan Pourzandi |
SSS | 4 |
| 2012 | Next-Generation Applications on Cellular Networks: Trends, Challenges, and SolutionsabstractApplications over cellular networks now range from operator–consumer applications (e.g., mobile television, voice-over-ip, video conferencing), peer-to-peer applications (e.g., instant messaging), machine-to-machine applications (e.g., data telemetry and automotive applications), mobile web services (e.g., music and video streaming), and social networking applications. The current approach for developing mobile applications appears to focus on utilizing template-based application-development kits provided by platform developers (e.g., Google's Android, Apple's iOS, or Nokia's Symbian) to capture application designs and install them on the runtime platforms through use of code generators tied to particular versions of the platform. It is still unclear as to how an application developer (or network operator) conceptualizes the features of a mobile application in a platform-independent way, identifies its utility and explores its impact on the user, or further refines the choice of technology, platform, and mobility/interactivity requirements. This paper attempts to offer some guidelines, based on recent research in the industry and academia in these areas, toward the design and development of successful mobile applications that can utilize the capabilities of the next generation of cellular networks. We provide an overview of the growing trends of the rich multimedia and real-time mobile applications, including the diversity of application types, their impact on the enterprise and consumer, their traffic volumes, and their load and communication patterns. In addition to the overall trend analysis, we also study the design choices that are to be made, and how they are realized, and also describe how the platforms (client and server) may be implemented. Additionally, we focus on mobile video applications according to their communication characteristics and their distinct demands on the cellular network. We also present an analysis of device and network application programming interfaces (API) that form the basic building blocks for efficient and secure mobile application development of the future. Nimish Radio, Ying Zhang 0022, Mallik Tatipamula, Vijay K. Madisetti |
Proc. IEEE | 2 |
| 2011 | ALIAS: scalable, decentralized label assignment for data centersabstractModern data centers can consist of hundreds of thousands of servers and millions of virtualized end hosts. Managing address assignment while simultaneously enabling scalable communication is a challenge in such an environment. We present ALIAS, an addressing and communication protocol that automates topology discovery and address assignment for the hierarchical topologies that underlie many data center network fabrics. Addresses assigned by ALIAS interoperate with a variety of scalable communication techniques. ALIAS is fully decentralized, scales to large network sizes, and dynamically recovers from arbitrary failures, without requiring modifications to hosts or to commodity switch hardware. We demonstrate through simulation that ALIAS quickly and correctly configures networks that support up to hundreds of thousands of hosts, even in the face of failures and erroneous cabling, and we show that ALIAS is a practical solution for auto-configuration with our NetFPGA testbed implementation. Meg Walraed-Sullivan, Radhika Niranjan Mysore, Malveeka Tewari, Ying Zhang 0022, Keith Marzullo, Amin Vahdat |
SoCC | 4 |
| 2011 | On Resilience of Split-Architecture NetworksabstractThe split architecture network assumes a logically centralized controller, which is physically separated from a large set of data plane forwarding switches. When the control plane becomes decoupled from the data plane, the requirement to the failure resilience and recovery mechanisms changes. In this work we investigate one of the most important practical issues in split architecture deployment, the placement of controllers in a given network. We first demonstrate that the location of controllers have high impact on the network resilience using a real network topology. Motivated by such observation, we propose a min-cut based controller placement algorithm and compare it with greedy based approach. Our simulation results show significant reliability improvements with an intelligent placement strategy. Our work is the first attempt on the resilience properties of a split architecture network. Ying Zhang 0022, Neda Beheshti, Mallik Tatipamula |
GLOBECOM | 1 |
| 2011 | A Comprehensive Long-Term Evaluation on BGP PerformanceabstractThe Border Gateway Protocol (BGP) is the de facto interdomain routing protocol on the Internet which controls the packet forwarding behavior on the data plane. It has significant impact on the well-being of the global Internet. Over the past ten years, there has been a large body of studies conducted on evaluating and improving the BGP performance. These studies develop tools using BGP data for identifying the Internet topology, AS relationships, and AS-level paths. More importantly BGP is the main data source for evaluating the Inter-domain routing performance and discovering routing anomalies such as prefix hijacking attacks. However, most of these studies focus on one or a few aspects of BGP in a short time period. Till today, the route monitoring system has been deployed for ten years and there has been a significant amount of criticisms on the bad performance of BGP. Our work is the first to critically examine and summarize BGP performance and its changes through time. We evaluate BGP from a diverse set of aspects ranging from routing diversity to convergence performance. We design a set of systematic statistical analysis to cope with the noise in data collection process. Due to the huge volume of data required for the analysis, we implement our evaluation system on top of the cloud computing platform from Amazon EC2. Our results provide a few insights on how to improve BGP and the Internet routing system in the future. Ying Zhang 0022, Mallik Tatipamula |
ICC | 1 |
| 2011 | Characterization and design of effective BGP AS-path prependingabstractThe AS path prepending approach in BGP is commonly used to perform inter-domain traffic engineering, such as inbound traffic load-balancing for multi-homed ASes. It artificially increases the length of the AS level path in BGP announcements by inserting its local AS number multiple times into outgoing EBGP announcement messages. In this work, we first present a comprehensive study on the characterization of Internet routing AS path prepending. We further propose an algorithm for computing the optimal padding strategies given multiple neighboring links. Our method considers the impact of AS relationship based local policies on ASPP's effectiveness. The algorithm can be used for three objectives, i.e., traffic load balancing, backup route provisioning, and bypassing a specific AS for security purposes, e.g., avoiding information censorship. We demonstrate the accuracy and effectiveness of our approach using real BGP data and traffic data from Abilene networks. Ying Zhang 0022, Mallik Tatipamula |
ICNP | 1 |
| 2010 | iSPY: Detecting IP Prefix Hijacking on My OwnabstractIP prefix hijacking remains a major threat to the security of the Internet routing system due to a lack of authoritative prefix ownership information. Despite many efforts in designing IP prefix hijack detection schemes, no existing design can satisfy all the critical requirements of a truly effective system: real-time, accurate, lightweight, easily and incrementally deployable, as well as robust in victim notification. In this paper, we present a novel approach that fulfills all these goals by monitoring network reachability from key external transit networks to one's own network through lightweight prefix-owner-based active probing. Using the prefix-owner's view of reachability, our detection system, iSPY, can differentiate between IP prefix hijacking and network failures based on the observation that hijacking is likely to result in topologically more diverse polluted networks and unreachability. Through detailed simulations of Internet routing, 25-day deployment in 88 autonomous systems (ASs) (108 prefixes), and experiments with hijacking events of our own prefix from multiple locations, we demonstrate that iSPY is accurate with false negative ratio below 0.45% and false positive ratio below 0.17%. Furthermore, iSPY is truly real-time; it can detect hijacking events within a few minutes. Zheng Zhang 0009, Ying Zhang 0022, Y. Charlie Hu, Z. Morley Mao, Randy Bush |
IEEE/ACM Trans. Netw. | 2 |
| 2009 | HC-BGP: A light-weight and flexible scheme for securing prefix ownershipabstractThe border gateway protocol (BGP) is a fundamental building block of the Internet infrastructure. However, due to the implicit trust assumption among networks, Internet routing remains quite vulnerable to various types of misconfiguration and attacks. Prefix hijacking is one such misbehavior where an attacker AS injects false routes to the Internet routing system that misleads victim's traffic to the attacker AS. Previous secure routing proposals, e.g., S-BGP, have relied on the global public key infrastructure (PKI), which creates deployment burdens. In this paper, we propose an efficient cryptographic mechanism, HC-BGP, using hash chains and regular public/private key pairs to ensure prefix ownership certificates. HC-BGP is computationally more efficient than previously proposed secure routing schemes, and it is also more flexible for supporting various traffic engineering goals. Our scheme can efficiently prevent common prefix hijacking attacks which announce routes with false origins, including both prefix and sub-prefix hijacking attacks. Ying Zhang 0022, Zheng Zhang 0009, Z. Morley Mao, Y. Charlie Hu |
DSN | 1 |
| 2009 | Detecting traffic differentiation in backbone ISPs with NetPoliceabstractTraffic differentiations are known to be found at the edge of the Internet in broadband ISPs and wireless carriers [13, 2]. The ability to detect traffic differentiations is essential for customers to develop effective strategies for improving their application performance. We build a system, called NetPolice, that enables detection of content- and routing-based differentiations in backbone ISPs. NetPolice is easy to deploy since it only relies on loss measurement launched from end hosts. The key challenges in building NetPolice include selecting an appropriate set of probing destinations and ensuring the robustness of detection results to measurement noise. Ying Zhang 0022, Z. Morley Mao, Ming Zhang 0005 |
Internet Measurement Conference | 1 |
| 2008 | Ascertaining the Reality of Network Neutrality Violation in Backbone ISPs
Ying Zhang 0022, Z. Morley Mao, Ming Zhang 0005 |
HotNets | 1 |
| 2008 | Effective Diagnosis of Routing Disruptions from End Systems
Ying Zhang 0022, Z. Morley Mao, Ming Zhang 0005 |
NSDI | 1 |
| 2008 | A Measurement Study of Internet Delay Asymmetry
Abhinav Pathak, Himabindu Pucha, Ying Zhang 0022, Y. Charlie Hu, Z. Morley Mao |
PAM | 3 |
| 2008 | Ispy: detecting ip prefix hijacking on my ownabstractIP prefix hijacking remains a major threat to the security of the Internet routing system due to a lack of authoritative prefix ownership information. Despite many efforts in designing IP prefix hijack detection schemes, no existing design can satisfy all the critical requirements of a truly effective system: real-time, accurate, light-weight, easily and incrementally deployable, as well as robust in victim notification. In this paper, we present a novel approach that fulfills all these goals by monitoring network reachability from key external transit networks to one's own network through lightweight prefix-owner-based active probing. Using the prefix-owner's view of reachability, our detection system, iSPY, can differentiate between IP prefix hijacking and network failures based on the observation that hijacking is likely to result in topologically more diverse polluted networks and unreachability. Through detailed simulations of Internet routing, 25-day deployment in 88 ASes (108 prefixes), and experiments with hijacking events of our own prefix from multiple locations, we demonstrate that iSPY is accurate with false negative ratio below 0.45% and false positive ratio below 0.17%. Furthermore, iSPY is truly real-time; it can detect hijacking events within a few minutes. Zheng Zhang 0009, Ying Zhang 0022, Y. Charlie Hu, Z. Morley Mao, Randy Bush |
SIGCOMM | 2 |
| 2007 | Internet routing resilience to failures: analysis and implicationsabstractInternet interdomain routing is policy-driven, and thus physical connectivity does not imply reachability. On average, routing on today's Internet works quite well, ensuring reachability for most networks and achieving reasonable performance across most paths. However, there is a serious lack of understanding of Internet routing resilience to significant but realistic failures such as those caused by the 911 event, the 2003 Northeast blackout, and the recent Taiwan earthquake in December 2006. In this paper, we systematically analyze how the current Internet routing system reacts to various types of failures by developing a realistic failure model, and then pinpoint reliability bottlenecks of the Internet. For validity of our simulation results, we generate topology graphs by addressing concerns over the incompleteness of topology and the inaccuracy of inferred AS relationships. By focusing on the impact of structural and policy properties, our analysis provides guidelines for future Internet design. The simulation tool we provide for analyzing routing resilience is also efficient to scale to Internet-size topologies. Jian Wu 0028, Ying Zhang 0022, Z. Morley Mao, Kang G. Shin |
CoNEXT | 2 |
| 2007 | Practical defenses against BGP prefix hijackingabstractPrefix hijacking, a misbehavior in which a misconfigured or malicious BGP router originates an IP prefix that the router does not own, is becoming an increasingly serious security problem on the Internet. In this paper, we conduct a first comprehensive study on incrementally deployable mitigation solutions against prefix hijacking. We first propose a novel reactive detection-assisted solution based on the idea of bogus route purging and valid route promotion. Our simulations based on realistic settings show that purging bogus routes at 20 highest-degree ASes reduces the polluted portion of the Internet by a random prefix hijack from 50% down to 24%, and adding promotion further reduces the remaining pollution by 33% ~ 57%, We prove that our proposed route purging and promotion scheme preserve the convergence properties of BGP regardless of the number of promoters. We are the first to demonstrate that detection systems based on a limited number of BGP feeds are subject to detection evasion by hijackers. Motivated the need for proactive defenses to complement reactive mitigation response, we evaluate customer route filtering, a best common practice among large ISPs today, and show its limited effectiveness. We also show the added benefits of combining route purging-promotion with customer route filtering. Zheng Zhang 0009, Ying Zhang 0022, Y. Charlie Hu, Z. Morley Mao |
CoNEXT | 2 |
| 2007 | A Firewall for Routers: Protecting against Routing MisbehaviorabstractIn this work, we present the novel idea of route normalization by correcting on the fly routing traffic on behalf of a local router to protect the local network from malicious and misconfigured routing updates. Analogous to traffic normalization for network intrusion detection systems, the proposed RouteNormalizer patches ambiguities and eliminates semantically incorrect routing updates to protect against routing protocol attacks. Furthermore, it serves the purpose of a router firewall by identifying resource-based attacks against routers. Upon detecting anomalous routing changes, it suggests local routing policy modifications to improve route selection decisions. Deploying a RouteNormalizer requires no modification to routers if desired using a transparent TCP proxy setup. In this paper, we present the detailed design of the RouteNormalizer and evaluate it using a prototype implementation based on empirical BGP routing updates. We validate its effectiveness by showing that many well-known routing problems from operator mailing lists are correctly identified. Ying Zhang 0022, Z. Morley Mao, Jia Wang 0001 |
DSN | 1 |
| 2007 | On the impact of route monitor selectionabstractSeveral route monitoring systems have been set up to help understand the Internet routing system. They operate by gathering real-time BGP updates from different networks. Many studies have relied on such data sources by assuming reasonably good coverage and thus representative visibility into the Internet routing system. However, different deployment strategies of route monitors directly impact the accuracy and generality of conclusions. Ying Zhang 0022, Zheng Zhang 0009, Z. Morley Mao, Y. Charlie Hu, Bruce M. Maggs |
Internet Measurement Conference | 1 |
| 2007 | A Framework for Measuring and Predicting the Impact of Routing ChangesabstractRouting dynamics heavily influence Internet data plane performance. Existing studies only narrowly focused on a few destinations and did not consider the predictability of the impact of routing changes on performance metrics such as reachability. In this work, we propose an efficient framework to capture coarse-grained but important performance degradation as a result of BGP routing events using light-weight probing. We deployed our framework across six vantage points for 11 weeks and found that the data plane experienced serious performance degradation in the form of reachability loss and forwarding loops following a significant fraction of updates affecting many destination prefixes and networks across all vantage points studied. Specifically, more than 39% of updates resulted in reachability loss, some lasting for more than 300 seconds, impacting more than 72% of probed prefixes and more than 35% of all the prefixes on the Internet. We identified that more than half of the prefixes have predictable routing behavior. Based on the stationarity of the correlation between routing changes and the data plane performance, we developed a model to accurately predict the severity of the impact due to routing changes. Such a model is directly helpful for making informed decisions for improved routing schemes such as overlay routing and backup path selection. Ying Zhang 0022, Z. Morley Mao, Jia Wang 0001 |
INFOCOM | 1 |
| 2007 | Low-Rate TCP-Targeted DoS Attack Disrupts Internet Routing
Ying Zhang 0022, Z. Morley Mao, Jia Wang 0001 |
NDSS | 1 |
| 2007 | Understanding network delay changes caused by routing eventsabstractNetwork delays and delay variations are two of the most important network performance metrics directly impacting real-time applications such as voice over IP and time-critical financial transactions. This importance is illustrated by past work on understanding the delay constancy of Internet paths and recent work on predicting network delays using virtual coordinate systems. Merely understanding currently observed delays is insufficient, as network performance can degrade not only due to traffic variability but also as a result of routing changes. Unfortunately this latter effect so far has been ignored in understanding and predicting delay related performance metrics of Internet paths. Our work is the first to address this short coming by systematically analyzing changes in network delays and jitter of a diverse and comprehensive set of Internet paths. Using empirical measurements, we illustrate that routing changes can result in roundtrip delay increase of converged paths by more than 1 second. Surprisingly, intradomain routing changes can also cause such large delay increase. Himabindu Pucha, Ying Zhang 0022, Z. Morley Mao, Y. Charlie Hu |
SIGMETRICS | 2 |