VLDB 2026 Research / reviewers in the wild / expert
Chuanxiong Guo
dblp:95/4962
· DBLP profile ↗
66ranked-venue papers
12as first author
9since 2021 · last 2026
0000-0002-0730-8468ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 49 · 12 first-author · 5 since 2021Systems, architecture and hardware · 9 · 2 since 2021Software engineering, systems software and programming languages · 6 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BROS: Efficient LLM Serving on Hybrid Real-time and Best-effort Requests
Borui Wan, Juntao Zhao 0002, Chenyu Jiang 0002, Chuanxiong Guo, Chuan Wu 0001 |
INFOCOM | 4 |
| 2023 | Lyra: Elastic Scheduling for Deep Learning ClustersabstractOrganizations often build separate training and inference clusters for deep learning, and use separate schedulers to manage them. This leads to problems for both: inference clusters have low utilization when the traffic load is low; training jobs often experience long queuing due to a lack of resources. We introduce Lyra, a new cluster scheduler to address these problems. Lyra introduces capacity loaning to loan idle inference servers for training jobs. It further exploits elastic scaling that scales a training job's resource allocation to better utilize loaned servers. Capacity loaning and elastic scaling create new challenges to cluster management. When the loaned servers need to be returned, we need to minimize job preemptions; when more GPUs become available, we need to allocate them to elastic jobs and minimize the job completion time (JCT). Lyra addresses these combinatorial problems with principled heuristics. It introduces the notion of server preemption cost, which it greedily reduces during server reclaiming. It further relies on the JCT reduction value defined for each additional worker of an elastic job to solve the scheduling problem as a multiple-choice knapsack problem. Prototype implementation on a 64-GPU testbed and large-scale simulation with 15-day traces of over 50,000 production jobs show that Lyra brings 1.53x and 1.48x reductions in average queuing time and JCT, and improves cluster usage by up to 25%. Jiamin Li 0002, Hong Xu 0001, Yibo Zhu 0001, Zherui Liu, Chuanxiong Guo, Cong Wang 0001 |
EuroSys | 5 |
| 2023 | BGL: GPU-Efficient GNN Training by Optimizing Graph Data I/O and Preprocessing
Tianfeng Liu, Yangrui Chen, Dan Li 0001, Chuan Wu 0001, Yibo Zhu 0001, Yanghua Peng, Hongzheng Chen, Chuanxiong Guo |
NSDI | 10 |
| 2023 | SRNIC: A Scalable Architecture for RDMA NICs
Zilong Wang 0007, Layong Luo, Qingsong Ning, Chaoliang Zeng, Wenxue Li 0004, Xinchen Wan, Xiongfei Geng, Tianhao Wang 0025, Weicheng Ling, Kejia Huo, Pingbo An, Kui Ji, Shideng Zhang, Ruiqing Feng, Kai Chen 0005, Chuanxiong Guo |
NSDI | 21 |
| 2023 | Predicting GPU Failures With High Precision Under Deep Learning WorkloadsabstractGraphics processing units (GPUs) are the de facto standard for processing deep learning (DL) tasks. In large-scale GPU clusters, GPU failures are inevitable and may cause severe consequences. For example, GPU failures disrupt distributed training, crash inference services, and result in service level agreement violations. In this paper, we study the problem of predicting GPU failures using machine learning (ML) models to mitigate their damages. Heting Liu, Cheng Tan 0005, Rongqiu Yang, Guohong Cao, Zherui Liu, Chuanxiong Guo |
SYSTOR | 7 |
| 2022 | Collie: Finding Performance Anomalies in RDMA Subsystems
Xinhao Kong, Yibo Zhu 0001, Huaping Zhou, Zhuo Jiang, Jianxi Ye, Chuanxiong Guo, Danyang Zhuo |
NSDI | 6 |
| 2022 | Tiara: A Scalable and Efficient Hardware Acceleration Architecture for Stateful Layer-4 Load Balancing
Chaoliang Zeng, Layong Luo, Zilong Wang 0007, Wenchen Han, Lebing Wan, Zhipeng Ding, Xiongfei Geng, Feng Ning, Kai Chen 0005, Chuanxiong Guo |
NSDI | 15 |
| 2022 | FAERY: An FPGA-accelerated Embedding-based Retrieval System
Chaoliang Zeng, Layong Luo, Qingsong Ning, Yaodong Han, Ding Tang, Zilong Wang 0007, Kai Chen 0005, Chuanxiong Guo |
OSDI | 9 |
| 2021 | AutoLRS: Automatic Learning-Rate Schedule by Bayesian Optimization on the Fly
Tianyi Zhou 0001, Liangyu Zhao, Yibo Zhu 0001, Chuanxiong Guo, Marco Canini, Arvind Krishnamurthy |
ICLR | 5 |
| 2020 | Elastic parameter server load distribution in deep learning clustersabstractIn distributed DNN training, parameter servers (PS) can become performance bottlenecks due to PS stragglers, caused by imbalanced parameter distribution, bandwidth contention, or computation interference. Few existing studies have investigated efficient parameter (aka load) distribution among PSs. We observe significant training inefficiency with the current parameter assignment in representative machine learning frameworks (e.g., MXNet, TensorFlow), and big potential for training acceleration with better PS load distribution. We design PSLD, a dynamic parameter server load distribution scheme, to mitigate PS straggler issues and accelerate distributed model training in the PS architecture. An exploitation-exploration method is carefully designed to scale in and out parameter servers and adjust parameter distribution among PSs on the go. We also design an elastic PS scaling module to carry out our scheme with little interruption to the training process. We implement our module on top of open-source PS architectures, including MXNet and BytePS. Testbed experiments show up to 2.86x speed-up in model training with PSLD, for different ML models under various straggler settings. Yangrui Chen, Yanghua Peng, Yixin Bao, Chuan Wu 0001, Yibo Zhu 0001, Chuanxiong Guo |
SoCC | 6 |
| 2020 | A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU Clusters
Yibo Zhu 0001, Chang Lan, Bairen Yi, Yong Cui 0001, Chuanxiong Guo |
OSDI | 6 |
| 2020 | Observing and Mitigating Micro-Burst Traffic in Data Center NetworksabstractMicro-burst traffic is not uncommon in data centers. It can cause packet dropping, which may result in serious performance degradation (e.g., Incast problem). However, current approaches to mitigate micro-burst is usually ad-hoc and not based on a principled understanding of the underlying behaviors. On the other hand, traditional studies focus on traffic burstiness in a single flow, while micro-burst traffic in the data centers could occur with highly fan-in communication pattern, and its dynamic behavior is still unclear. To this end, in this paper, we re-examine the micro-burst traffic in typical data center scenarios. We find that the evolution of micro-burst is determined by both TCP's self-clocking mechanism and congestion control algorithm. Besides, dynamic behaviors of micro-burst under various scenarios can all be described by the time derivative of queue length evolution.Our observations also implicate that conventional solutions like absorbing and pacing are ineffective to mitigate micro-burst traffic.Instead, senders need to rapidly respond to some explicit signals of the queue buildup caused by the micro-burst traffic rather than independently and ineffectually pacing themselves in isolation. Inspired by the findings and insights from experimental observations, we propose Micro-burst-Aware Transport Control Protocol (MATCP), which leverages characteristic behaviors of micro-burst traffic derived from the time derivative of the queue occupancy. MATCP can suppress the sharp queue length increment by over 2x and reduce the tail query completion time by up to 84.4%. Danfeng Shan, Fengyuan Ren, Peng Cheng 0005, Ran Shu 0001, Chuanxiong Guo |
IEEE/ACM Trans. Netw. | 5 |
| 2019 | Tiresias: A GPU Cluster Manager for Distributed Deep Learning
Juncheng Gu, Mosharaf Chowdhury, Kang G. Shin, Yibo Zhu 0001, Myeongjae Jeon, Junjie Qian, Hongqiang Harry Liu, Chuanxiong Guo |
NSDI | 8 |
| 2019 | FreeFlow: Software-based Virtual RDMA Networking for Containerized Clouds
Daehyeok Kim, Tianlong Yu, Hongqiang Harry Liu, Yibo Zhu 0001, Jitendra Padhye, Shachar Raindel, Chuanxiong Guo, Vyas Sekar, Srinivasan Seshan |
NSDI | 7 |
| 2019 | NetBouncer: Active Device and Link Failure Localization in Data Center Networks
Cheng Tan 0005, Ze Jin, Chuanxiong Guo, Tianrong Zhang, Karl Deng, Dongming Bi |
NSDI | 3 |
| 2019 | A generic communication scheduler for distributed DNN training accelerationabstractWe present ByteScheduler, a generic communication scheduler for distributed DNN training acceleration. ByteScheduler is based on our principled analysis that partitioning and rearranging the tensor transmissions can result in optimal results in theory and good performance in real-world even with scheduling overhead. To make ByteScheduler work generally for various DNN training frameworks, we introduce a unified abstraction and a Dependency Proxy mechanism to enable communication scheduling without breaking the original dependencies in framework engines. We further introduce a Bayesian Optimization approach to auto-tune tensor partition size and other parameters for different training models under various networking conditions. ByteScheduler now supports TensorFlow, PyTorch, and MXNet without modifying their source code, and works well with both Parameter Server (PS) and all-reduce architectures for gradient synchronization, using either TCP or RDMA. Our experiments show that ByteScheduler accelerates training with all experimented system configurations and DNN models, by up to 196% (or 2.96X of original speed). Yanghua Peng, Yibo Zhu 0001, Yangrui Chen, Yixin Bao, Bairen Yi, Chang Lan, Chuan Wu 0001, Chuanxiong Guo |
SOSP | 8 |
| 2019 | Tagger: Practical PFC Deadlock Prevention in Data Center NetworksabstractRemote direct memory access over converged Ethernet deployments is vulnerable to deadlocks induced by priority flow control. Prior solutions for deadlock prevention either require significant changes to routing protocols or require excessive buffers in the switches. In this paper, we propose Tagger, a scheme for deadlock prevention. It does not require any changes to the routing protocol and needs only modest buffers. Tagger is based on the insight that given a set of expected lossless routes, a simple tagging scheme can be developed to ensure that no deadlock will occur under any failure conditions. Packets that do not travel on these lossless routes may be dropped under extreme conditions. We design such a scheme, prove that it prevents deadlock, and implement it efficiently on commodity hardware. Shuihai Hu, Yibo Zhu 0001, Peng Cheng 0005, Chuanxiong Guo, Jitendra Padhye, Kai Chen 0005 |
IEEE/ACM Trans. Netw. | 4 |
| 2018 | Optimus: an efficient dynamic resource scheduler for deep learning clustersabstractDeep learning workloads are common in today's production clusters due to the proliferation of deep learning driven AI services (e.g., speech recognition, machine translation). A deep learning training job is resource-intensive and time-consuming. Efficient resource scheduling is the key to the maximal performance of a deep learning cluster. Existing cluster schedulers are largely not tailored to deep learning jobs, and typically specifying a fixed amount of resources for each job, prohibiting high resource efficiency and job performance. This paper proposes Optimus, a customized job scheduler for deep learning clusters, which minimizes job training time based on online resource-performance models. Optimus uses online fitting to predict model convergence during training, and sets up performance models to accurately estimate training speed as a function of allocated resources in each job. Based on the models, a simple yet effective method is designed and used for dynamically allocating resources and placing deep learning tasks to minimize job completion time. We implement Optimus on top of Kubernetes, a cluster manager for container orchestration, and experiment on a deep learning cluster with 7 CPU servers and 6 GPU servers, running 9 training jobs using the MXNet framework. Results show that Optimus outperforms representative cluster schedulers by about 139% and 63% in terms of job completion time and makespan, respectively. Yanghua Peng, Yixin Bao, Yangrui Chen, Chuan Wu 0001, Chuanxiong Guo |
EuroSys | 5 |
| 2018 | Micro-Burst in Data Centers: Observations, Analysis, and MitigationsabstractMicro-burst traffic is not uncommon in data centers. It can cause packet dropping, which results in serious performance degradation (e.g., Incast problem). However, current solutions that attempt to suppress micro-burst traffic are extrinsic and ad hoc, since they lack the comprehensive and essential understanding of micro-burst's root cause and dynamic behavior. On the other hand, traditional studies focus on traffic burstiness in a single flow, while in data centers micro-burst traffic could occur with highly fan-in communication pattern, and its dynamic behavior is still unclear. To this end, in this paper, we re-examine the microburst traffic in typical data center scenarios. We find that evolution of micro-burst is determined by both TCP's self-clocking mechanism and bottleneck link. Besides, dynamic behaviors of micro-burst under various scenarios can all be described by the slope of queue length evolution. Our observations also implicate that conventional solutions like absorbing and pacing are ineffective to mitigate micro-burst traffic. Instead, senders need to slow down as soon as possible. Inspired by the findings and insights from experimental observations, we propose S-ECN policy, which is an ECN marking policy leveraging the slope of queue length evolution. Transport protocols utilizing S-ECN policy can suppress the sharp queue length increment by over 2×, and reduce the average query completion time by ~12-27%. Danfeng Shan, Fengyuan Ren, Peng Cheng 0005, Ran Shu 0001, Chuanxiong Guo |
ICNP | 5 |
| 2018 | Deepview: Virtual Disk Failure Diagnosis and Pattern Detection for Azure
Qiao Zhang 0001, Chuanxiong Guo, Yingnong Dang, Nick Swanson, Xinsheng Yang, Randolph Yao, Murali Chintalapati, Arvind Krishnamurthy, Thomas E. Anderson |
NSDI | 3 |
| 2018 | Capturing and Enhancing In Situ System Observability for Failure Detection
Peng Huang 0005, Chuanxiong Guo, Jacob R. Lorch, Lidong Zhou, Yingnong Dang |
OSDI | 2 |
| 2018 | TerseCades: Efficient Data Compression in Stream Processing
Gennady Pekhimenko, Chuanxiong Guo, Myeongjae Jeon, Peng Huang 0005, Lidong Zhou |
USENIX ATC | 2 |
| 2017 | Tagger: Practical PFC Deadlock Prevention in Data Center NetworksabstractRemote Direct Memory Access over Converged Ethernet (RoCE) deployments are vulnerable to deadlocks induced by Priority Flow Control (PFC). Prior solutions for deadlock prevention either require signi.cant changes to routing protocols, or require excessive bu.ers in the switches. In this paper, we propose Tagger, a scheme for deadlock prevention. It does not require any changes to the routing protocol, and needs only modest bu.ers. Tagger is based on the insight that given a set of expected lossless routes, a simple tagging scheme can be developed to ensure that no deadlock will occur under any failure conditions. Packets that do not travel on these lossless routes may be dropped under extreme conditions. We design such a scheme, prove that it prevents deadlock and implement it e.ciently on commodity hardware. Shuihai Hu, Peng Cheng 0005, Chuanxiong Guo, Jitendra Padhye, Kai Chen 0005 |
CoNEXT | 4 |
| 2017 | Gray Failure: The Achilles' Heel of Cloud-Scale SystemsabstractCloud scale provides the vast resources necessary to replace failed components, but this is useful only if those failures can be detected. For this reason, the major availability breakdowns and performance anomalies we see in cloud environments tend to be caused by subtle underlying faults, i.e., gray failure rather than fail-stop failure. In this paper, we discuss our experiences with gray failure in production cloud-scale systems to show its broad scope and consequences. We also argue that a key feature of gray failure is differential observability: that the system's failure detectors may not notice problems even when applications are afflicted by them. This realization leads us to believe that, to best deal with them, we should focus on bridging the gap between different components' perceptions of what constitutes failure. Peng Huang 0005, Chuanxiong Guo, Lidong Zhou, Jacob R. Lorch, Yingnong Dang, Murali Chintalapati, Randolph Yao |
HotOS | 2 |
| 2017 | Virtualized Network Coding Functions on the InternetabstractNetwork coding is a fundamental tool that enables higher network capacity and lower complexity in routing algorithms, by encouraging the mixing of information flows in the middle of a network. Implementing network coding in the core Internet is subject to practical concerns, since Internet routers are often overwhelmed by packet forwarding tasks, leaving little processing capacity for coding operations. Inspired by the recent paradigm of network function virtualization, we propose implementing network coding as a new network function, and deploying such coding functions in geo-distributed cloud data centers, to practically enable network coding on the Internet. We target multicast sessions (including unicast flows as special cases), strategically deploy relay nodes (network coding functions) in selected data centers between senders and receivers, and embrace high bandwidth efficiency brought by network coding with dynamic coding function deployment. We design and implement the network coding function on typical virtual machines, featuring efficient packet processing. We propose an efficient algorithm for coding function deployment, scaling in and out, in the presence of system dynamics. Real-world implementation on Amazon EC2 and Linode demonstrates significant throughput improvement and higher robustness of multicast via coding functions as well as efficiency of the dynamic deployment and scaling algorithm. Linquan Zhang, Shangqi Lai, Chuan Wu 0001, Zongpeng Li, Chuanxiong Guo |
ICDCS | 5 |
| 2017 | deTector: a Topology-aware Monitoring System for Data Center Networks
Yanghua Peng, Chuan Wu 0001, Chuanxiong Guo, Chengchen Hu, Zongpeng Li |
USENIX ATC | 4 |
| 2017 | Orchestrating Bulk Data Transfers across Geo-Distributed DatacentersabstractAs it has become the norm for cloud providers to host multiple datacenters around the globe, significant demands exist for inter-datacenter data transfers in large volumes, e.g., migration of big data. A challenge arises on how to schedule the bulk data transfers at different urgency levels, in order to fully utilize the available inter-datacenter bandwidth. The Software Defined Networking (SDN) paradigm has emerged recently which decouples the control plane from the data paths, enabling potential global optimization of data routing in a network. This paper aims to design a dynamic, highly efficient bulk data transfer service in a geo-distributed datacenter system, and engineer its design and solution algorithms closely within an SDN architecture. We model data transfer demands as delay tolerant migration requests with different finishing deadlines. Thanks to the flexibility provided by SDN, we enable dynamic, optimal routing of distinct chunks within each bulk data transfer (instead of treating each transfer as an infinite flow), which can be temporarily stored at intermediate datacenters to mitigate bandwidth contention with more urgent transfers. An optimal chunk routing optimization model is formulated to solve for the best chunk transfer schedules over time. To derive the optimal schedules in an online fashion, three algorithms are discussed, namely a bandwidth-reserving algorithm, a dynamically-adjusting algorithm, and a future-demand-friendly algorithm, targeting at different levels of optimality and scalability. We build an SDN system based on the Beacon platform and OpenFlow APIs, and carefully engineer our bulk data transfer algorithms in the system. Extensive real-world experiments are carried out to compare the three algorithms as well as those from the existing literature, in terms of routing optimality, computational delay and overhead. Yu Wu 0010, Chuan Wu 0001, Chuanxiong Guo, Zongpeng Li, Francis C. M. Lau 0001 |
IEEE Trans. Cloud Comput. | 4 |
| 2017 | CubicRing: Exploiting Network Proximity for Distributed In-Memory Key-Value StoreabstractIn-memory storage has the benefits of low I/O latency and high I/O throughput. Fast failure recovery is crucial for large-scale in-memory storage systems, bringing network-related challenges, including false detection due to transient network problems, traffic congestion during the recovery, and top-of-rack switch failures. In order to achieve fast failure recovery, in this paper, we present CubicRing, a distributed structure for cube-based networks, which exploits network proximity to restrict failure detection and recovery within the smallest possible one-hop range. We leverage the CubicRing structure to address the aforementioned challenges and design a network-aware in-memory key-value store called MemCube. In a 64-node 10GbE testbed, MemCube recovers 48 GB of data for a single server failure in 3.1 s. The 14 recovery servers achieve 123.9 Gb/s aggregate recovery throughput, which is 88.5% of the ideal aggregate bandwidth and several times faster than RAMCloud with the same configurations. Yiming Zhang 0003, Dongsheng Li 0001, Chuanxiong Guo, Yongqiang Xiong, Xicheng Lu |
IEEE/ACM Trans. Netw. | 3 |
| 2016 | Deadlocks in Datacenter Networks: Why Do They Form, and How to Avoid ThemabstractDriven by the need for ultra-low latency, high throughput and low CPU overhead, Remote Direct Memory Access (RDMA) is being deployed by many cloud providers. To deploy RDMA in Ethernet networks, Priority-based Flow Control (PFC) must be used. PFC, however, makes Ethernet networks prone to deadlocks. Prior work on deadlock avoidance has focused on {\em necessary} condition for deadlock formation, which leads to rather onerous and expensive solutions for deadlock avoidance. In this paper, we investigate {\em sufficient} conditions for deadlock formation, conjecturing that avoiding {\em sufficient} conditions might be less onerous. Shuihai Hu, Peng Cheng 0005, Chuanxiong Guo, Jitendra Padhye, Kai Chen 0005 |
HotNets | 4 |
| 2016 | RDMA over Commodity Ethernet at ScaleabstractOver the past one and half years, we have been using RDMA over commodity Ethernet (RoCEv2) to support some of Microsoft's highly-reliable, latency-sensitive services. This paper describes the challenges we encountered during the process and the solutions we devised to address them. In order to scale RoCEv2 beyond VLAN, we have designed a DSCP-based priority flow-control (PFC) mechanism to ensure large-scale deployment. We have addressed the safety challenges brought by PFC-induced deadlock (yes, it happened!), RDMA transport livelock, and the NIC PFC pause frame storm problem. We have also built the monitoring and management systems to make sure RDMA works as expected. Our experiences show that the safety and scalability issues of running RoCEv2 at scale can all be addressed, and RDMA can replace TCP for intra data center communications and achieve low latency, low CPU overhead, and high throughput. Chuanxiong Guo, Zhong Deng, Gaurav Soni, Jianxi Ye, Jitendra Padhye, Marina Lipshteyn |
SIGCOMM | 1 |
| 2016 | Explicit Path Control in Commodity Data Centers: Design and ApplicationsabstractMany data center network DCN applications require explicit routing path control over the underlying topologies. In this paper, we present XPath, a simple, practical and readily-deployable way to implement explicit path control, using existing commodity switches. At its core, XPath explicitly identifies an end-to-end path with a path ID and leverages a two-step compression algorithm to pre-install all the desired paths into IP TCAM tables of commodity switches. Our evaluation and implementation show that XPath scales to large DCNs and is readily-deployable. Furthermore, on our testbed, we integrate XPath into four applications to showcase its utility. Shuihai Hu, Kai Chen 0005, Wei Bai 0001, Chang Lan, Hao Wang 0022, Chuanxiong Guo |
IEEE/ACM Trans. Netw. | 8 |
| 2015 | UniDrive: Synergize Multiple Consumer Cloud Storage ServicesabstractConsumer cloud storage (CCS) services have become popular among users for storing and synchronizing files via apps installed on their devices. A single CCS, however, has intrinsic limitations on networking performance, service reliability, and data security. To overcome these limitations, we present UniDrive, a CCS app that synergizes multiple CCSs (multi-cloud) by using only few simple public RESTful Web APIs. UniDrive follows a server-less, client-centric design, in which synchronization logic is purely implemented at client devices and all communication is conveyed through file upload and download operations. Strong consistency of the metadata is guaranteed via a quorum-based distributed mutual-exclusive lock mechanism. UniDrive improves reliability and security by judiciously distributing erasure coded files across multiple CCSs. To boost networking performance, UniDrive leverages all available clouds to maximize parallel transfer opportunities, but the key insight behind is the concept of data block over-provisioning and dynamic scheduling. This suite of techniques masks the diversified and varying network conditions of the underlying clouds, and exploits more the faster clouds via a simple yet effective in-channel probing scheme. Extensive experimental results on the global Amazon EC2 platform and a real-world trial by 272 users confirmed significantly superior and consistent sync performance of UniDrive over any single CCS. Haowen Tang, Fangming Liu, Guobin Shen, Chuanxiong Guo |
Middleware | 5 |
| 2015 | Explicit Path Control in Commodity Data Centers: Design and Applications
Shuihai Hu, Kai Chen 0005, Wei Bai 0001, Chang Lan, Hao Wang 0022, Chuanxiong Guo |
NSDI | 8 |
| 2015 | CubicRing: Enabling One-Hop Failure Detection and Recovery for Distributed In-Memory Storage Systems
Yiming Zhang 0003, Chuanxiong Guo, Dongsheng Li 0001, Rui Chu, Yongqiang Xiong |
NSDI | 2 |
| 2015 | Pingmesh: A Large-Scale System for Data Center Network Latency Measurement and AnalysisabstractCan we get network latency between any two servers at any time in large-scale data center networks? The collected latency data can then be used to address a series of challenges: telling if an application perceived latency issue is caused by the network or not, defining and tracking network service level agreement (SLA), and automatic network troubleshooting. We have developed the Pingmesh system for large-scale data center network latency measurement and analysis to answer the above question affirmatively. Pingmesh has been running in Microsoft data centers for more than four years, and it collects tens of terabytes of latency data per day. Pingmesh is widely used by not only network software developers and engineers, but also application and service developers and operators. Chuanxiong Guo, Yingnong Dang, Ray Huang, David A. Maltz, Vin Wang, Varugis Kurien |
SIGCOMM | 1 |
| 2015 | Congestion Control for Large-Scale RDMA DeploymentsabstractModern datacenter applications demand high throughput (40Gbps) and ultra-low latency (< 10 μs per hop) from the network, with low CPU overhead. Standard TCP/IP stacks cannot meet these requirements, but Remote Direct Memory Access (RDMA) can. On IP-routed datacenter networks, RDMA is deployed using RoCEv2 protocol, which relies on Priority-based Flow Control (PFC) to enable a drop-free network. However, PFC can lead to poor application performance due to problems like head-of-line blocking and unfairness. To alleviates these problems, we introduce DCQCN, an end-to-end congestion control scheme for RoCEv2. To optimize DCQCN performance, we build a fluid model, and provide guidelines for tuning switch buffer thresholds, and other protocol parameters. Using a 3-tier Clos network testbed, we show that DCQCN dramatically improves throughput and fairness of RoCEv2 RDMA traffic. DCQCN is implemented in Mellanox NICs, and is being deployed in Microsoft's datacenters. Yibo Zhu 0001, Haggai Eran, Daniel Firestone, Chuanxiong Guo, Marina Lipshteyn, Yehonatan Liron, Jitendra Padhye, Shachar Raindel, Mohamad Haj Yahia, Ming Zhang 0005 |
SIGCOMM | 4 |
| 2013 | Per-packet load-balanced, low-latency routing for clos-based data center networksabstractClos-based networks including Fat-tree and VL2 are being built in data centers, but existing per-flow based routing causes low network utilization and long latency tail. In this paper, by studying the structural properties of Fat-tree and VL2, we propose a per-packet round-robin based routing algorithm called Digit-Reversal Bouncing (DRB). DRB achieves perfect packet interleaving. Our analysis and simulations show that, compared with random-based load-balancing algorithms, DRB results in smaller and bounded queues even when traffic load approaches 100%, and it uses smaller re-sequencing buffer for absorbing out-of-order packet arrivals. Our implementation demonstrates that our design can be readily implemented with commodity switches. Experiments on our testbed, a Fat-tree with 54 servers, confirm our analysis and simulations, and further show that our design handles network failures in 1-2 seconds and has the desirable graceful performance degradation property. Jiaxin Cao, Pengkun Yang, Chuanxiong Guo, Guohan Lu, Yixin Zheng, Yongqiang Xiong, David A. Maltz |
CoNEXT | 4 |
| 2013 | PACE: Policy-Aware Application Cloud EmbeddingabstractThe emergence of new capabilities such as virtualization and elastic (private or public) cloud computing infrastructures has made it possible to deploy multiple applications, on demand, on the same cloud infrastructure. A major challenge to achieve this possibility, however, is that modern applications are typically distributed, structured systems that include not only computational and storage entities, but also policy entities (e.g., load balancers, firewalls, intrusion prevention boxes). Deploying applications on a cloud infrastructure without the policy entities may introduce substantial policy violations and/or security holes. In this paper, we present PACE: the first systematic framework for Policy-Aware Application Cloud Embedding. We precisely define the policy-aware, cloud application embedding problem, study its complexity and introduce simple, efficient, online primal-dual algorithms to embed applications in cloud data centers. We conduct evaluations using data from a real, large campus network and a realistic data center topology to evaluate the feasibility and performance of PACE. We show that deployment in a cloud without considering in-network policies may lead to a large number of policy violations (e.g., using tree routing as a way to enforce in-network policies may observe up to 91% policy violations). We also show that our embedding algorithms are very efficient by comparing with a good online fractional embedding algorithm. Li Erran Li, Vahid Liaghat, Mohammad Hajiaghayi, Dan Li 0001, Gordon T. Wilfong, Yang Richard Yang, Chuanxiong Guo |
INFOCOM | 8 |
| 2013 | Moving big data to the cloudabstractCloud computing, rapidly emerging as a new computation paradigm, provides agile and scalable resource access in a utility-like fashion, especially for the processing of big data. An important open issue here is how to efficiently move the data, from different geographical locations over time, into a cloud for effective processing. The de facto approach of hard drive shipping is not flexible, nor secure. This work studies timely, cost-minimizing upload of massive, dynamically-generated, geodispersed data into the cloud, for processing using a MapReducelike framework. Targeting at a cloud encompassing disparate data centers, we model a cost-minimizing data migration problem, and propose two online algorithms, for optimizing at any given time the choice of the data center for data aggregation and processing, as well as the routes for transmitting data there. The first is an online lazy migration (OLM) algorithm achieving a competitive ratio of as low as 2.55, under typical system settings. The second is a randomized fixed horizon control (RFHC) algorithm achieving a competitive ratio of 1+ 1/l+λ κ/λ with a lookahead window of l, where κ and λ are system parameters of similar magnitude. Linquan Zhang, Chuan Wu 0001, Zongpeng Li, Chuanxiong Guo, Minghua Chen 0001, Francis C. M. Lau 0001 |
INFOCOM | 4 |
| 2013 | Datacast: A Scalable and Efficient Reliable Group Data Delivery Service for Data CentersabstractReliable Group Data Delivery (RGDD) is a pervasive traffic pattern in data centers. In an RGDD group, a sender needs to reliably deliver a copy of data to all the receivers. Existing solutions either do not scale due to the large number of RGDD groups (e.g., IP multicast) or cannot efficiently use network bandwidth (e.g., end-host overlays). Motivated by recent advances on data center network topology designs (multiple edge-disjoint Steiner trees for RGDD) and innovations on network devices (practical in-network packet caching), we propose Datacast for RGDD. Datacast explores two design spaces: 1) Datacast uses multiple edge-disjoint Steiner trees for data delivery acceleration. 2) Datacast leverages in-network packet caching and introduces a simple soft-state based congestion control algorithm to address the scalability and efficiency issues of RGDD. Our analysis reveals that Datacast congestion control works well with small cache sizes (e.g., 125KB) and causes few duplicate data transmissions (e.g., 1.19%). Both simulations and experiments confirm our theoretical analysis. We also use experiments to compare the performance of Datacast and BitTorrent. In a BCube(4, 1) with 1Gbps links, we use both Datacast and BitTorrent to transmit 4GB data. The link stress of Datacast is 1.01, while it is 1.39 for BitTorrent. By using two Steiner trees, Datacast finishes the transmission in 16.9s, while BitTorrent uses 52s. Jiaxin Cao, Chuanxiong Guo, Guohan Lu, Yongqiang Xiong, Yixin Zheng, Yongguang Zhang, Yibo Zhu 0001, Chen Chen 0019, Ye Tian 0004 |
IEEE J. Sel. Areas Commun. | 2 |
| 2013 | Moving Big Data to The Cloud: An Online Cost-Minimizing ApproachabstractCloud computing, rapidly emerging as a new computation paradigm, provides agile and scalable resource access in a utility-like fashion, especially for the processing of big data. An important open issue here is to efficiently move the data, from different geographical locations over time, into a cloud for effective processing. The de facto approach of hard drive shipping is not flexible or secure. This work studies timely, cost-minimizing upload of massive, dynamically-generated, geo-dispersed data into the cloud, for processing using a MapReduce-like framework. Targeting at a cloud encompassing disparate data centers, we model a cost-minimizing data migration problem, and propose two online algorithms: an online lazy migration (OLM) algorithm and a randomized fixed horizon control (RFHC) algorithm , for optimizing at any given time the choice of the data center for data aggregation and processing, as well as the routes for transmitting data there. Careful comparisons among these online and offline algorithms in realistic settings are conducted through extensive experiments, which demonstrate close-to-offline-optimum performance of the online algorithms. Linquan Zhang, Chuan Wu 0001, Zongpeng Li, Chuanxiong Guo, Minghua Chen 0001, Francis C. M. Lau 0001 |
IEEE J. Sel. Areas Commun. | 4 |
| 2013 | ICTCP: Incast Congestion Control for TCP in Data-Center NetworksabstractTransport Control Protocol (TCP) incast congestion happens in high-bandwidth and low-latency networks when multiple synchronized servers send data to the same receiver in parallel. For many important data-center applications such as MapReduce and Search, this many-to-one traffic pattern is common. Hence TCP incast congestion may severely degrade their performances, e.g., by increasing response time. In this paper, we study TCP incast in detail by focusing on the relationships between TCP throughput, round-trip time (RTT), and receive window. Unlike previous approaches, which mitigate the impact of TCP incast congestion by using a fine-grained timeout value, our idea is to design an Incast congestion Control for TCP (ICTCP) scheme on the receiver side. In particular, our method adjusts the TCP receive window proactively before packet loss occurs. The implementation and experiments in our testbed demonstrate that we achieve almost zero timeouts and high goodput for TCP incast. Zhenqian Feng, Chuanxiong Guo, Yongguang Zhang |
IEEE/ACM Trans. Netw. | 3 |
| 2013 | IP-Geolocation Mapping for Moderately Connected Internet RegionsabstractMost IP-geolocation mapping schemes [14], [16], [17], [18] take delay-measurement approach, based on the assumption of a strong correlation between networking delay and geographical distance between the targeted client and the landmarks. In this paper, however, we investigate a large region of moderately connected Internet and find the delay-distance correlation is weak. But we discover a more probable rule - with high probability the shortest delay comes from the closest distance. Based on this closest-shortest rule, we develop a simple and novel IP-geolocation mapping scheme for moderately connected Internet regions, called GeoGet. In GeoGet, we take a large number of webservers as passive landmarks and map a targeted client to the geolocation of the landmark that has the shortest delay. We further use JavaScript at targeted clients to generate HTTP/Get probing for delay measurement. To control the measurement cost, we adopt a multistep probing method to refine the geolocation of a targeted client, finally to city level. The evaluation results show that when probing about 100 landmarks, GeoGet correctly maps 35.4 percent clients to city level, which outperforms current schemes such as GeoLim [16] and GeoPing [14] by 270 and 239 percent, respectively, and the median error distance in GeoGet is around 120 km, outperforming GeoLim and GeoPing by 37 and 70 percent, respectively. Dan Li 0001, Chuanxiong Guo, Yunxin Liu 0001, Zhi-Li Zhang, Yongguang Zhang |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2012 | Datacast: a scalable and efficient reliable group data delivery service for data centersabstractReliable Group Data Delivery (RGDD) is a pervasive traffic pattern in data centers. In an RGDD group, a sender needs to reliably deliver a copy of data to all the receivers. Existing solutions either do not scale due to the large number of RGDD groups (e.g., IP multicast) or cannot efficiently use network bandwidth (e.g., end-host overlays). Jiaxin Cao, Chuanxiong Guo, Guohan Lu, Yongqiang Xiong, Yixin Zheng, Yongguang Zhang, Yibo Zhu 0001, Chen Chen 0019 |
CoNEXT | 2 |
| 2012 | Tuning ECN for data center networksabstractThere have been some serious concerns about the TCP performance in data center networks, including the long completion time of short TCP flows in competition with long TCP flows, and the congestion due to TCP incast. In this paper, we show that a properly tuned instant queue length based Explicit Congestion Notification (ECN) at the intermediate switches can alleviate both problems. Compared with previous work, our approach is appealing as it can be supported on current commodity switches with a simple parameter setting and it does not need any modification on ECN protocol at the end servers. Furthermore, we have observed a dilemma in which a higher ECN threshold leads to higher throughput for long flows whereas a lower threshold leads to more senders on incast under buffer pressure. We address this problem with a switch modification only scheme - dequeue marking, for further tuning the instant queue length based ECN to achieve optimal incast performance and long flow throughput with a single threshold value. Our experimental study demonstrates that dequeue marking is effective for increasing the maximum incast senders close to the performance limit of ECN, achieving a gain anywhere from 16% to 140%. Jiabo Ju, Guohan Lu, Chuanxiong Guo, Yongqiang Xiong, Yongguang Zhang |
CoNEXT | 4 |
| 2012 | DAC: Generic and Automatic Address Configuration for Data Center NetworksabstractData center networks encode locality and topology information into their server and switch addresses for performance and routing purposes. For this reason, the traditional address configuration protocols such as DHCP require a huge amount of manual input, leaving them error-prone. In this paper, we present DAC, a generic and automatic Data center Address Configuration system. With an automatically generated blueprint that defines the connections of servers and switches labeled by logical IDs, e.g., IP addresses, DAC first learns the physical topology labeled by device IDs, e.g., MAC addresses. Then, at the core of DAC is its device-to-logical ID mapping and malfunction detection. DAC makes an innovation in abstracting the device-to-logical ID mapping to the graph isomorphism problem and solves it with low time complexity by leveraging the attributes of data center network topologies. Its malfunction detection scheme detects errors such as device and link failures and miswirings, including the most difficult case where miswirings do not cause any node degree change. We have evaluated DAC via simulation, implementation, and experiments. Our simulation results show that DAC can accurately find all the hardest-to-detect malfunctions and can autoconfigure a large data center with 3.8 million devices in 46 s. In our implementation, we successfully autoconfigure a small 64-server BCube network within 300 ms and show that DAC is a viable solution for data center autoconfiguration. Kai Chen 0005, Chuanxiong Guo, Zhenqian Feng, Yan Chen 0004, Songwu Lu, Wenfei Wu |
IEEE/ACM Trans. Netw. | 2 |
| 2011 | RDCM: Reliable data center multicastabstractMulticast benefits data center group communication in both saving network traffic and improving application throughput. The SLA (Service Level Agreement) of cloud service requires the computation correctness of distributed applications, translating to the requirement of reliable Multicast delivery. In this paper we present RDCM, a novel reliable Multicast approach for data center network. The key idea of RDCM is to minimize the impact of packet loss on the Multicast performance, by leveraging the rich link resource in data centers. A Multicast-tree-aware backup overlay is purposely built on group members for peer-to-peer packet repair. Riding on Unicast, packet repair not only achieves complete repair isolation, but also has high probability to bypass the pathological links in the Multicast tree where packet loss occurs. The backup overlay is organized in such a way that it causes little individual repair burden, control overhead, as well as overall repair traffic. We have implemented RDCM as a user-level library on Windows platform. The experiments on our test bed show that RDCM handles packet loss without obvious throughput degradation during high-speed data transmission. Dan Li 0001, Mingwei Xu 0001, Ming-Chen Zhao, Chuanxiong Guo, Yongguang Zhang, Min-You Wu |
INFOCOM | 4 |
| 2011 | ServerSwitch: A Programmable and High Performance Platform for Data Center Networks
Guohan Lu, Chuanxiong Guo, Tong Yuan, Yongqiang Xiong, Yongguang Zhang |
NSDI | 2 |
| 2011 | Scalable and cost-effective interconnection of data-center servers using dual server portsabstractThe goal of data-center networking is to interconnect a large number of server machines with low equipment cost while providing high network capacity and high bisection width. It is well understood that the current practice where servers are connected by a tree hierarchy of network switches cannot meet these requirements. In this paper, we explore a new server-interconnection structure. We observe that the commodity server machines used in today's data centers usually come with two built-in Ethernet ports, one for network connection and the other left for backup purposes. We believe that if both ports are actively used in network connections, we can build a scalable, cost-effective interconnection structure without either the expensive higher-level large switches or any additional hardware on servers. We design such a networking structure called FiConn. Although the server node degree is only 2 in this structure, we have proven that FiConn is highly scalable to encompass hundreds of thousands of servers with low diameter and high bisection width. We have developed a low-overhead traffic-aware routing mechanism to improve effective link utilization based on dynamic traffic state. We have also proposed how to incrementally deploy FiConn. Dan Li 0001, Chuanxiong Guo, Kun Tan 0001, Yongguang Zhang, Songwu Lu |
IEEE/ACM Trans. Netw. | 2 |
| 2010 | SecondNet: a data center network virtualization architecture with bandwidth guaranteesabstractIn this paper, we propose virtual data center (VDC) as the unit of resource allocation for multiple tenants in the cloud. VDCs are more desirable than physical data centers because the resources allocated to VDCs can be rapidly adjusted as tenants' needs change. To enable the VDC abstraction, we design a data center network virtualization architecture called SecondNet. SecondNet achieves scalability by distributing all the virtual-to-physical mapping, routing, and bandwidth reservation state in server hypervisors. Its port-switching based source routing (PSSR) further makes SecondNet applicable to arbitrary network topologies using commodity servers and switches. SecondNet introduces a centralized VDC allocation algorithm for bandwidth guaranteed virtual to physical mapping. Simulations demonstrate that our VDC allocation achieves high network utilization and low time complexity. Our implementation and experiments show that we can build SecondNet on top of various network topologies, and SecondNet provides bandwidth guarantee and elasticity, as designed. Chuanxiong Guo, Guohan Lu, Helen J. Wang, Chao Kong, Wenfei Wu, Yongguang Zhang |
CoNEXT | 1 |
| 2010 | ICTCP: Incast Congestion Control for TCP in data center networksabstractTCP incast congestion happens in high-bandwidth and low-latency networks, when multiple synchronized servers send data to a same receiver in parallel [15]. For many important data center applications such as MapReduce[5] and Search, this many-to-one traffic pattern is common. Hence TCP in-cast congestion may severely degrade their performances, e.g., by increasing response time. Zhenqian Feng, Chuanxiong Guo, Yongguang Zhang |
CoNEXT | 3 |
| 2010 | Generic and automatic address configuration for data center networksabstractData center networks encode locality and topology information into their server and switch addresses for performance and routing purposes. For this reason, the traditional address configuration protocols such as DHCP require huge amount of manual input, leaving them error-prone.In this paper, we present DAC, a generic and automatic Data center Address Configuration system. With an automatically generated blueprint which defines the connections of servers and switches labeled by logical IDs, e.g., IP addresses, DAC first learns the physical topology labeled by device IDs, e.g., MAC addresses. Then at the core of DAC is its device-to-logical ID mapping and malfunction detection. DAC makes an innovation in abstracting the device-to-logical ID mapping to the graph isomorphism problem, and solves it with low time-complexity by leveraging the attributes of data center network topologies. Its malfunction detection scheme detects errors such as device and link failures and miswirings, including the most difficult case where miswirings do not cause any node degree change.We have evaluated DAC via simulation, implementation and experiments. Our simulation results show that DAC can accurately find all the hardest-to-detect malfunctions and can autoconfigure a large data center with 3.8 million devices in 46 seconds. In our implementation, we successfully autoconfigure a small 64-server BCube network within 300 milliseconds and show that DAC is a viable solution for data center autoconfiguration. Kai Chen 0005, Chuanxiong Guo, Zhenqian Feng, Yan Chen 0004, Songwu Lu, Wenfei Wu |
SIGCOMM | 2 |
| 2009 | MDCube: a high performance network structure for modular data center interconnectionabstractShipping-container-based data centers have been introduced as building blocks for constructing mega-data centers. However, it is a challenge on how to interconnect those containers together with reasonable cost and cabling complexity, due to the fact that a mega-data center can have hundreds or even thousands of containers and the aggregate bandwidth among containers can easily reach tera-bit per second. As a new inner-container server-centric network architecture, BCube [9] interconnects thousands of servers inside a container and provides high bandwidth support for typical traffic patterns. It naturally serves as a building block for mega-data center. Guohan Lu, Dan Li 0001, Chuanxiong Guo, Yongguang Zhang |
CoNEXT | 4 |
| 2009 | Mining the Web and the Internet for Accurate IP Address GeolocationsabstractIn this paper, we present Structon, a novel approach that uses Web mining together with inference and IP traceroute to geolocate IP addresses with significantly better accuracy than existing automated approaches. Structon is composed of three ideas which we realize in three corresponding steps. First, we extract geolocation information of Web server IP addresses from Web pages. Second, we devise heuristic algorithms to improve both the accuracy and the coverage of the IP geolocation database using these Web server IP addresses and their geolocations as input. Third, for those segments that are not covered in the first two steps, we use IP traceroute to identify the access routers of those segments. When the location of the access router is known, we can deduce the location of the associated segment since it is co-located together with the access router. By mining 500-million Web pages collected in China in 2006 (11 percent of the total Web pages in China at that time), we are able to identify the geolocations for 103 million IP addresses. This represents nearly 88 percent IP addresses allocated to China in March 2008. Structon is 87.4 percent accurate at city granularity and up to 93.5 percent accurate at province level. We also used 10 day Windows Live client log to evaluate our client IP addresses coverage: Structon identified geolocations of 98.9 percent of client IP addresses. Chuanxiong Guo, Yunxin Liu 0001, Wenchao Shen, Helen J. Wang, Yongguang Zhang |
INFOCOM | 1 |
| 2009 | FiConn: Using Backup Port for Server Interconnection in Data CentersabstractThe goal of data center networking is to interconnect a large number of server machines with low equipment cost, high and balanced network capacity, and robustness to link/server faults. It is well understood that, the current practice where servers are connected by a tree hierarchy of network switches cannot meet these requirements (Fares et al., 2008 and Guo et al., 2008). In this paper, we explore a new server-interconnection structure. We observe that the commodity server machines used in today's data centers usually come with two built-in Ethernet ports, one for network connection and the other left for backup purpose. We believe that, if both ports are actively used in network connections, we can build a low-cost interconnection structure without the expensive higher-level large switches. Our new network design, called FiConn, utilizes both ports and only the low-end commodity switches to form a scalable and highly effective structure. Although the server node degree is only two in this structure, we have proven that FiConn is highly scalable to encompass hundreds of thousands of servers with low diameter and high bisection width. The routing mechanism in FiConn balances different levels of links. We have further developed a low-overhead traffic-aware routing mechanism to improve effective link utilization based on dynamic traffic state. Simulation results have demonstrated that the routing mechanisms indeed achieve high networking throughput. Dan Li 0001, Chuanxiong Guo, Kun Tan 0001, Songwu Lu |
INFOCOM | 2 |
| 2009 | A scalable micro wireless interconnect structure for CMPsabstractThis paper describes an unconventional way to apply wireless networking in emerging technologies. It makes the case for using a two-tier hybrid wireless/wired architecture to interconnect hundreds to thousands of cores in chip multiprocessors (CMPs), where current interconnect technologies face severe scaling limitations in excessive latency, long wiring, and complex layout. We propose a recursive wireless interconnect structure called the WCube that features a single transmit antenna and multiple receive antennas at each micro wireless router and offers scalable performance in terms of latency and connectivity. We show the feasibility to build miniature on-chip antennas, and simple transmitters and receivers that operate at 100 − 500 GHz sub-terahertz frequency bands. We also devise new two-tier wormhole based routing algorithms that are deadlock free and ensure a minimum-latency route on a 1000core on-chip interconnect network. Our simulations show that our protocol suite can reduce the observed latency by 20 % to 45%, and consumes power that is comparable to or less than current 2-D wiredmeshdesigns. Suk-Bok Lee, Sai-Wang Tam, Ioannis Pefkianakis, Songwu Lu, Mau-Chung Frank Chang, Chuanxiong Guo, Glenn Reinman, Chunyi Peng 0001, Mishali Naik, Lixia Zhang 0001, Jason Cong |
MobiCom | 6 |
| 2009 | BCube: a high performance, server-centric network architecture for modular data centersabstractThis paper presents BCube, a new network architecture specifically designed for shipping-container based, modular data centers. At the core of the BCube architecture is its server-centric network structure, where servers with multiple network ports connect to multiple layers of COTS (commodity off-the-shelf) mini-switches. Servers act as not only end hosts, but also relay nodes for each other. BCube supports various bandwidth-intensive applications by speeding-up one-to-one, one-to-several, and one-to-all traffic patterns, and by providing high network capacity for all-to-all traffic. Chuanxiong Guo, Guohan Lu, Dan Li 0001, Yunfeng Shi, Chen Tian 0001, Yongguang Zhang, Songwu Lu |
SIGCOMM | 1 |
| 2008 | Improved Smoothed Round Robin Schedulers for High-Speed Packet NetworksabstractThe smoothed round robin (SRR) packet scheduler is attractive for use in high-speed networks due to its very low time complexity, but it is not suitable for realtime applications since it cannot provide tight delay bound. In this paper, we present two improved algorithms based on SRR, namely SRRcand SRR#, which are based on novel matrix transform techniques. By transforming the irregular Weight Matrix of SRR into triangular and diagonal ones, SRR+and SRR#are able to evenly interleave flows based on their reserved rates even in skewed weight distributions. SRR+and SRR#provide bounded delay, whereas are still of low space and time complexities and are simple to implement in high-speed networks. The properties of SRR+and SRR#are addressed in detail through analysis and simulations. SRR+and SRR#, together with SRR and the recently developed G-3 scheduler [9] form a full spectrum of schedulers that provide tradeoffs among delay, time complexity, and space complexity. Chuanxiong Guo |
INFOCOM | 1 |
| 2008 | Dcell: a scalable and fault-tolerant network structure for data centers
Chuanxiong Guo, Kun Tan 0001, Lei Shi 0002, Yongguang Zhang, Songwu Lu |
SIGCOMM | 1 |
| 2007 | G-3: An O(1) Time Complexity Packet Scheduler That Provides Bounded End-to-End DelayabstractIn this paper, we present anO(1) time-complexity packet scheduling algorithm which we call G-3 that provides bounded end-to-end delay for fixed size packet networks. G-3 is built over two round-robin schedulers SRR (Chuanxiong Guo, 2004) and RRR (Garg and Xiaoqiang Chen, 1999) and several novel data structures. In G-3, bounded delay is provided by evenly distributing the binary coded weight of a flow into a square weight matrix (SWM) and several perfect weighted binary trees (PWBTs). In order to achieveO(1) time complexity, the SWM matrix is further spread by a weight spread sequence (WSS) and each PWBT tree is spread by a corresponding time-slot sequence (TSS), respectively. G-3 then performs packet scheduling by sequential scanning the WSS and TSS sequences. G-3 can be implemented in high-speed packet networks to provide bandwidth guarantee, fairness, and bounded delay due to itsO(1) time complexity. Chuanxiong Guo |
INFOCOM | 1 |
| 2007 | Generic Application-Level Protocol Analyzer and its Language
Nikita Borisov, David Brumley, Helen J. Wang, John Dunagan, Pallavi Joshi, Chuanxiong Guo |
NDSS | 6 |
| 2005 | End-system-based mobility support in IPv6abstractNumerous mobility solutions have been proposed in the past, but none of them have been widely deployed today. To address the deployment difficulty in previous work, we propose an end-system-based mobility solution for IPv6 (EMIPv6). In our design, we adhere to the end-to-end principle (Saltzer et al., 1984) by directly performing connection maintenance and data packet delivery between the two communicating hosts. And we leverage distributed hash table-based peer-to-peer (P2P) systems to carry out self-organized, scalable, and robust name lookup for mobile hosts. When the mobility messages such as address updates cannot be delivered directly between the end hosts (e.g., due to firewalls, network address translators, or simultaneous movement), we propose a distributed subscription/notification (S/N) service on top of the previously introduced P2P overlay to deliver them. This leads to a complete end-system solution with small handoff latency and efficient packet delivery. Our simulation results showed that our scheme achieves small name resolution latency by considering host heterogeneity into the design. We have implemented EMIPv6-based end-systems. The experiments in our testbed demonstrated that a complete end-system-based mobility solution is technically feasible and is easy to deploy in the real world without the need of introducing new network components. Chuanxiong Guo, Kun Tan 0001, Qian Zhang 0001, Jingmin Song, Junfeng Zhou, Christian Huitema, Wenwu Zhu 0001 |
IEEE J. Sel. Areas Commun. | 1 |
| 2004 | Shield: vulnerability-driven network filters for preventing known vulnerability exploitsabstractSoftware patching has not been effective as a first-line defense against large-scale worm attacks, even when patches have long been available for their corresponding vulnerabilities. Generally, people have been reluctant to patch their systems immediately, because patches are perceived to be unreliable and disruptive to apply. To address this problem, we propose a first-line worm defense in the network stack, using shields -- vulnerability-specific, exploit-generic network filters installed in end systems once a vulnerability is discovered, but before a patch is applied. These filters examine the incoming or outgoing traffic of vulnerable applications, and correct traffic that exploits vulnerabilities. Shields are less disruptive to install and uninstall, easier to test for bad side effects, and hence more reliable than traditional software patches. Further, shields are resilient to polymorphic or metamorphic variations of exploits [43].In this paper, we show that this concept is feasible by describing a prototype Shield framework implementation that filters traffic above the transport layer. We have designed a safe and restrictive language to describe vulnerabilities as partial state machines of the vulnerable application. The expressiveness of the language has been verified by encoding the signatures of several known vulnerabilites. Our evaluation provides evidence of Shield's low false positive rate and small impact on application throughput. An examination of a sample set of known vulnerabilities suggests that Shield could be used to prevent exploitation of a substantial fraction of the most dangerous ones. Helen J. Wang, Chuanxiong Guo, Daniel R. Simon, Alf Zugenmaier |
SIGCOMM | 2 |
| 2004 | A seamless and proactive end-to-end mobility solution for roaming across heterogeneous wireless networksabstractRoaming across heterogeneous wireless networks such as wireless wide area network (WWAN) and wireless local area network (WLAN) poses considerable challenges, as it is usually difficult to maintain the existing connections and guarantee the necessary quality-of-service. This paper proposes a novel seamless and proactive end-to-end mobility management system, which can maintain the connections based on the end-to-end principle by incorporating an intelligent network status detection mechanism. The proposed system consists of two components, connection manager (CM) and virtual connectivity (VC). The CM, by using novel media access control-layer and physical-layer sensing techniques, can obtain accurate network condition, while at the same time reducing the unnecessary handoff and ping-pong effect. The VC can make mobility transparent to applications without additional network-layer infrastructure support using a local connection translation, and can handle mobility well in the network address translator and simultaneous movement cases using a subscription/notification service. The proposed system enjoys several unique advantages: 1) capable of reacting to roaming events proactively and accurately; 2) maintaining the connection's continuity with small handoff delay; and 3) being a unified end-to-end approach for both IPv4 and IPv6 networks. We have built a prototype system and performed experiments to demonstrate the advantages of the proposed system. Chuanxiong Guo, Zihua Guo, Qian Zhang 0001, Wenwu Zhu 0001 |
IEEE J. Sel. Areas Commun. | 1 |
| 2004 | SRR: an O(1) time-complexity packet scheduler for flows in multiservice packet networksabstractWe present a novel fair queueing scheme, which we call Smoothed Round Robin (SRR). Ordinary round-robin schedulers are well known for the burstiness of their scheduling output. In order to overcome this problem, SRR codes the weights of the flows into binary vectors to form a Weight Matrix and then uses a Weight Spread Sequence (WSS), which is specially designed to distribute the output more evenly, to schedule packets by scanning a Weight Matrix. By using the WSS and the Weight Matrix, SRR emulates the Generalized Processor Sharing (GPS) well. SRR possesses better short-term fairness and scheduling delay properties in comparison with various existing round-robin schedulers. At the same time, SRR preserves O(1) time complexity by avoiding the time-stamp maintenance employed in various fair queueing schedulers. Simulation and implementation experiments show that SRR provides good mean end-to-end delay for soft real-time services. SRR can be implemented in high-speed networks to provide quality of service due to its simplicity and low time complexity. Chuanxiong Guo |
IEEE/ACM Trans. Netw. | 1 |
| 2001 | SRR: An O(1) time complexity packet scheduler for flows in multi-service packet networksabstractIn this paper, we present a novel fair queueing scheme, which we call Smoothed Round Robin (SRR). Ordinary round robin schedulers are well known for their burstiness in the scheduling output. In order to overcome this problem, SRR codes the weights of the flows into binary vectors to form a Weight Matrix, then uses a Weight Spread Sequence (WSS), which is specially designed to distribute the output more evenly, to schedule packets by scanning the Weight Matrix. By using the WSS and the Weight Matrix, SRR can emulate the Generalized Processor Sharing (GPS) well. It possesses better short-term fairness and schedule delay properties in comparison with various round robin schedulers. At the same time, it preserves O(1) time complexity by avoiding the time-stamp maintenance employed in various Fair Queueing schedulers. Simulation and implementation experiments show that SRR can provide good average end-to-end delay for soft real-time services. SRR can also be implemented in highspeed networks to provide QoS for its simplicity and low time complexity. Chuanxiong Guo |
SIGCOMM | 1 |