EDBT 2026 Demo / reviewers in the wild / expert
Yibo Zhu 0001
dblp:65/8854-1
· DBLP profile ↗
54ranked-venue papers
5as first author
25since 2021 · last 2026
0000-0002-9113-2660ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 32 · 5 first-author · 10 since 2021Systems, architecture and hardware · 10 · 9 since 2021Software engineering, systems software and programming languages · 7 · 3 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic Sparsity in Large-Scale Video DiT TrainingabstractDiffusion Transformers (DiTs) have shown remarkable performance in generating high-quality videos. However, the quadratic complexity of 3D full attention remains a bottleneck in scaling DiT training, especially with high-definition, lengthy videos, where it can consume up to 95% of processing time and demand specialized context parallelism. Xin Tan 0004, Yuetao Chen, Xing Chen 0009, Kun Yan 0004, Nan Duan 0001, Yibo Zhu 0001, Daxin Jiang, Hong Xu 0001 |
ASPLOS (1) | 7 |
| 2026 | DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved PipelineabstractLarge multimodal models (LMMs) have demonstrated excellent capabilities in both understanding and generation tasks with various modalities. While these models can accept flexible combinations of input data, their training efficiency suffers from two major issues: pipeline stage imbalance caused by heterogeneous model architectures, and training data dynamicity stemming from the diversity of multimodal data. Zhenliang Xue, Hanpeng Hu, Xing Chen 0009, Yixin Song 0003, Zeyu Mi, Yibo Zhu 0001, Daxin Jiang, Yubin Xia, Haibo Chen 0001 |
ASPLOS (2) | 7 |
| 2026 | Dynamic Compute and Network Orchestration for Disaggregated RLabstractDisaggregating the generation and training stages in RL is widely adopted to scale LLM post-training. There are two critical challenges here. First, the generation stage often becomes a bottleneck due to dynamic workload shifts and severe execution imbalances. Second, the decoupled stages result in diverse and dynamic network traffic patterns that strain the conventional static fabric. Xin Tan 0004, Yicheng Feng, Yu Zhou 0008, Yibo Zhu 0001, Hong Xu 0001 |
SIGCOMM | 5 |
| 2025 | Optimizing RLHF Training for Large Language Models with Stage Fusion
Yinmin Zhong, Bingyang Wu, Changyi Wan, Hanpeng Hu, Ranchen Ming, Yibo Zhu 0001, Xin Jin 0008 |
NSDI | 10 |
| 2025 | InfiniteHBD: Building Datacenter-Scale High-Bandwidth Domain for LLM with Optical Circuit Switching TransceiversabstractScaling Large Language Model (LLM) training relies on multidimensional parallelism, where High-Bandwidth Domains (HBDs) are critical for communication-intensive parallelism like Tensor Parallelism. However, existing HBD architectures face fundamental limitations in scalability, cost, and fault resiliency: switch-centric HBDs (e.g., NVL-72) incur prohibitive scaling costs, while GPU-centric HBDs (e.g., TPUv3/Dojo) suffer from severe fault propagation. Switch-GPU hybrid HBDs (e.g., TPUv4) take a middle-ground approach, but the fault explosion radius remains large. Chenchen Shou, Guyue Liu, Hao Nie, Huaiyu Meng, Yu Zhou 0008, Wenqing Lv, Yelong Xu, Yuanwei Lu, Yanbo Yu, Yichen Shen 0001, Yibo Zhu 0001, Daxin Jiang |
SIGCOMM | 13 |
| 2025 | DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language ModelsabstractMultimodal large language models (LLMs) empower LLMs to ingest inputs and generate outputs in multiple forms, such as text, image, and audio. However, the integration of multiple modalities introduces heterogeneity in both the model and training data, creating unique systems challenges. Yinmin Zhong, Hanpeng Hu, Jianjian Sun, Zheng Ge, Yibo Zhu 0001, Daxin Jiang, Xin Jin 0008 |
SIGCOMM | 7 |
| 2024 | CDMPP: A Device-Model Agnostic Framework for Latency Prediction of Tensor ProgramsabstractDeep Neural Networks (DNNs) have shown excellent performance in a wide range of machine learning applications. Knowing the latency of running a DNN model or tensor program on a specific device is useful in various tasks, such as DNN graph- or tensor-level optimization and device selection. Considering the large space of DNN models and devices that impedes direct profiling of all combinations, recent efforts focus on building a predictor to model the performance of DNN models on different devices. However, none of the existing attempts have achieved a cost model that can accurately predict the performance of various tensor programs while supporting both training and inference accelerators. We propose CDMPP, an efficient tensor program latency prediction framework for both cross-model and cross-device prediction. We design an informative but efficient representation of tensor programs, called compact ASTs, and a pre-order-based positional encoding method, to capture the internal structure of tensor programs. We develop a domain-adaption-inspired method to learn domain-invariant representations and devise a KMeans-based sampling algorithm, for the predictor to learn from different domains (i.e., different DNN operators and devices). Our extensive experiments on a diverse range of DNN models and devices demonstrate that CDMPP significantly outperforms state-of-the-art baselines with 14.03% and 10.85% prediction error for cross-model and cross-device prediction, respectively, and one order of magnitude higher training efficiency. The implementation and the expanded dataset are available at https://github.com/joapolarbear/cdmpp. Hanpeng Hu, Junwei Su, Juntao Zhao 0002, Yanghua Peng, Yibo Zhu 0001, Haibin Lin, Chuan Wu 0001 |
EuroSys | 5 |
| 2024 | QSync: Quantization-Minimized Synchronous Distributed Training Across Hybrid DevicesabstractA number of production deep learning clusters have attempted to explore inference hardware for DNN training, at the off-peak serving hours with many inference GPUs idling. Conducting DNN training with a combination of heterogeneous training and inference GPUs, known as hybrid device training, presents considerable challenges due to disparities in compute capability and significant differences in memory capacity. We propose QSync, a training system that enables efficient synchronous data-parallel DNN training over hybrid devices by strategically exploiting quantized operators. According to each device’s available resource capacity, QSync selects a quantization-minimized setting for operators in the distributed DNN training graph, minimizing model accuracy degradation but keeping the training efficiency brought by quantization. We carefully design a predictor with a bi-directional mixed-precision indicator to reflect the sensitivity of DNN layers on fixed-point and floating-point low-precision operators, a replayer with a neighborhood-aware cost mapper to accurately estimate the latency of distributed hybrid mixed-precision training, and then an allocator that efficiently synchronizes workers with minimized model accuracy degradation. QSync bridges the computational graph on PyTorch to an optimized backend for quantization kernel performance and flexible support for various GPU architectures. Extensive experiments show that QSync’s predictor can accurately simulate distributed mixed-precision training with < 5% error, with a consistent 0.27 − 1.03% accuracy improvement over the from-scratch training tasks compared to uniform precision. Juntao Zhao 0002, Borui Wan, Yanghua Peng, Haibin Lin, Yibo Zhu 0001, Chuan Wu 0001 |
IPDPS | 5 |
| 2024 | DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
Yinmin Zhong, Junda Chen, Jianbo Hu, Yibo Zhu 0001, Xuanzhe Liu, Xin Jin 0008, Hao Zhang 0025 |
OSDI | 5 |
| 2024 | MuxFlow: efficient GPU sharing in production-level clusters with more than 10000 GPUs
Xuanzhe Liu, Shufan Liu, Xiang Li 0067, Yibo Zhu 0001, Xin Liu 0086, Xin Jin 0008 |
Sci. China Inf. Sci. | 5 |
| 2024 | DistMind: Efficient Resource Disaggregation for Deep Learning WorkloadsabstractDeep learning (DL) systems suffer from low resource utilization due to 1) monolithic server model that tightly couples compute and memory; and 2) limited sharing between different inference applications, and across inference and training, because of strict service level objectives (SLOs). To address this problem, we present, a disaggregated DL system that enables efficient multiplexing of DL applications with near-optimal resource utilization. decouples compute from host memory, and exposes the abstractions of a GPU pool and a memory pool, each of which can be independently provisioned. The key challenge is to dynamically allocate GPU resources to different applications based on their real-time demands while meeting strict SLOs. We tackle this challenge by exploiting the power of high-speed 100 Gbps networks, and design three-stage pipelining, cache-aware load balancing, and DNN-aware sharding mechanisms based on the characteristics of DL workloads, to achieve millisecond-scale application loading overhead and improve system efficiency. We have implemented a prototype of and integrated it with PyTorch. Experimental results on AWS EC2 show that achieves near 100% resource utilization, and compared with NVIDIA MPS and Ray, improves the throughput by up to 279% and reduces the inference latency by up to 94%. Xin Jin 0008, Zhihao Bai, Zhen Zhang 0063, Yibo Zhu 0001, Yinmin Zhong, Xuanzhe Liu |
IEEE/ACM Trans. Netw. | 4 |
| 2023 | Lyra: Elastic Scheduling for Deep Learning ClustersabstractOrganizations often build separate training and inference clusters for deep learning, and use separate schedulers to manage them. This leads to problems for both: inference clusters have low utilization when the traffic load is low; training jobs often experience long queuing due to a lack of resources. We introduce Lyra, a new cluster scheduler to address these problems. Lyra introduces capacity loaning to loan idle inference servers for training jobs. It further exploits elastic scaling that scales a training job's resource allocation to better utilize loaned servers. Capacity loaning and elastic scaling create new challenges to cluster management. When the loaned servers need to be returned, we need to minimize job preemptions; when more GPUs become available, we need to allocate them to elastic jobs and minimize the job completion time (JCT). Lyra addresses these combinatorial problems with principled heuristics. It introduces the notion of server preemption cost, which it greedily reduces during server reclaiming. It further relies on the JCT reduction value defined for each additional worker of an elastic job to solve the scheduling problem as a multiple-choice knapsack problem. Prototype implementation on a 64-GPU testbed and large-scale simulation with 15-day traces of over 50,000 production jobs show that Lyra brings 1.53x and 1.48x reductions in average queuing time and JCT, and improves cluster usage by up to 25%. Jiamin Li 0002, Hong Xu 0001, Yibo Zhu 0001, Zherui Liu, Chuanxiong Guo, Cong Wang 0001 |
EuroSys | 3 |
| 2023 | Hi-Speed DNN Training with Espresso: Unleashing the Full Potential of Gradient Compression with Near-Optimal Usage StrategiesabstractGradient compression (GC) is a promising approach to addressing the communication bottleneck in distributed deep learning (DDL). It saves the communication time, but also incurs additional computation overheads. The training throughput of compression-enabled DDL is determined by the compression strategy, including whether to compress each tensor, the type of compute resources (e.g., CPUs or GPUs) for compression, the communication schemes for compressed tensor, and so on. However, it is challenging to find the optimal compression strategy for applying GC to DDL because of the intricate interactions among tensors. To fully unleash the benefits of GC, two questions must be addressed: 1) How to express any compression strategies and the corresponding interactions among tensors of any DDL training job? 2) How to quickly select a near-optimal compression strategy? Haibin Lin, Yibo Zhu 0001, T. S. Eugene Ng |
EuroSys | 3 |
| 2023 | ByteTransformer: A High-Performance Transformer Boosted for Variable-Length InputsabstractTransformers have become keystone models in natural language processing over the past decade. They have achieved great popularity in deep learning applications, but the increasing sizes of the parameter spaces required by transformer models generate a commensurate need to accelerate performance. Natural language processing problems are also routinely faced with variable-length sequences, as word counts commonly vary among sentences. Existing deep learning frameworks pad variable-length sequences to a maximal length, which adds significant memory and computational overhead. In this paper, we present ByteTransformer, a high-performance transformer boosted for variable-length inputs. We propose a padding-free algorithm that liberates the entire transformer from redundant computations on zero padded tokens. In addition to algorithmic-level optimization, we provide architecture-aware optimizations for transformer functional modules, especially the performance-critical algorithm Multi-Head Attention (MHA). Experimental results on an NVIDIA A100 GPU with variable-length sequence inputs validate that our fused MHA outperforms PyTorch by 6.13x. The end-to-end performance of ByteTransformer for a forward BERT transformer surpasses state-of-the-art transformer frameworks, such as PyTorch JIT, TensorFlow XLA, Tencent TurboTransformer, Microsoft DeepSpeed-Inference and NVIDIA FasterTransformer, by 87%, 131%, 138%, 74% and 55%, respectively. We also demonstrate the general applicability of our optimization methods to other BERT-like models, including ALBERT, DistilBERT, and DeBERTa. Chengquan Jiang, Leyuan Wang, Xiaoying Jia 0001, Zizhong Chen, Xin Liu 0086, Yibo Zhu 0001 |
IPDPS | 8 |
| 2023 | BGL: GPU-Efficient GNN Training by Optimizing Graph Data I/O and Preprocessing
Tianfeng Liu, Yangrui Chen, Dan Li 0001, Chuan Wu 0001, Yibo Zhu 0001, Yanghua Peng, Hongzheng Chen, Chuanxiong Guo |
NSDI | 5 |
| 2023 | Accelerating Distributed MoE Training and Inference with Lina
Jiamin Li 0002, Yibo Zhu 0001, Cong Wang 0001, Hong Xu 0001 |
USENIX ATC | 3 |
| 2023 | Discrete Cosin TransFormer: Image Modeling From Frequency DomainabstractIn this paper, we propose Discrete Cosin TransFormer (DCFormer) that directly learn semantics from DCT-based frequency domain representation. We first show that transformer-based networks are able to learn semantics directly from frequency domain representation based on discrete cosine transform (DCT) without compromising the performance. To achieve the desired efficiency-effectiveness trade-off, we then leverage an input information compression on its frequency domain representation, which highlights the visually significant signals inspired by JPEG compression. We explore different frequency domain downsampling strategies and show that it is possible to preserve the semantic meaningful information by strategically dropping the high-frequency components. The proposed DCFormer is tested on various downstream tasks including image classification, object detection and instance segmentation, and achieves state-of-the-art comparable performance with less FLOPs, and outperforms the commonly used backbone (e.g. SWIN) at similar FLOPs. Our ablation results also show that the proposed method generalizes well on different transformer backbones. Yanyi Zhang, Hanlin Lu, Yibo Zhu 0001 |
WACV | 5 |
| 2023 | SP-GNN: Learning structure and position information from graphs
Yangrui Chen, Jiaxuan You, Yanghua Peng, Chuan Wu 0001, Yibo Zhu 0001 |
Neural Networks | 7 |
| 2022 | SAPipe: Staleness-Aware Pipeline for Data Parallel DNN TrainingabstractData parallelism across multiple machines is widely adopted for accelerating distributed deep learning, but it is hard to achieve linear speedup due to the heavy communication. In this paper, we propose SAPipe, a performant system that pushes the training speed of data parallelism to its fullest extent. By introducing partial staleness, the communication overlaps the computation with minimal staleness in SAPipe. To mitigate additional problems incurred by staleness, SAPipe adopts staleness compensation techniques including weight prediction and delay compensation with provably lower error bounds. Additionally, SAPipe presents an algorithm-system co-design with runtime optimization to minimize system overhead for the staleness training pipeline and staleness compensation. We have implemented SAPipe in the BytePS framework, compatible to both TensorFlow and PyTorch. Our experiments show that SAPipe achieves up to 157% speedups over BytePS (non-stale), and outperforms PipeSGD in accuracy by up to 13.7%. Yangrui Chen, Juncheng Gu, Yanghua Peng, Haibin Lin, Chuan Wu 0001, Yibo Zhu 0001 |
NeurIPS | 8 |
| 2022 | Collie: Finding Performance Anomalies in RDMA Subsystems
Xinhao Kong, Yibo Zhu 0001, Huaping Zhou, Zhuo Jiang, Jianxi Ye, Chuanxiong Guo, Danyang Zhuo |
NSDI | 2 |
| 2022 | Multi-resource interleaving for deep learning trainingabstractTraining Deep Learning (DL) model requires multiple resource types, including CPUs, GPUs, storage IO, and network IO. Advancements in DL have produced a wide spectrum of models that have diverse usage patterns on different resource types. Existing DL schedulers focus on only GPU allocation, while missing the opportunity of packing jobs along multiple resource types. Yuanqiang Liu, Yanghua Peng, Yibo Zhu 0001, Xuanzhe Liu, Xin Jin 0008 |
SIGCOMM | 4 |
| 2022 | Congestion Control for Cross-Datacenter NetworksabstractGeographically distributed applications hosted on cloud are becoming prevalent. They run oncross-datacenter networkthat consists of multiple data center networks (DCNs) connected by a wide area network (WAN). Such a cross-DC network poses significant challenges in transport design because the DCN and WAN segments have vastly distinct characteristics (e.g., buffer depths, RTTs). In this paper, we find that existing DCN or WAN transport reacting to ECN or delay alone do not (and cannot be extended to) work well for such an environment. The key reason is that neither of the signals, by itself only, can simultaneously capture the location and degree of congestion, mainly due to the discrepancies between DCN and WAN. Motivated by this, we present the design and implementation of GEMINI that strategically integrates both ECN and delay signals for cross-DC congestion control. To achieve low latency, GEMINI bounds the inter-DC latency with delay signal and prevents the intra-DC packet loss with ECN. To maintain high throughput, GEMINI modulates the window dynamics and maintains low buffer occupancy utilizing both congestion signals. GEMINI is implemented in Linux kernel and evaluated by extensive testbed experiments. Results show that GEMINI achieves up to 53%, 31%, 76% and 2% reduction of small flow average completion times, and up to 34%, 39%, 9% and 58% reduction of large flow average completion times compared to TCP Cubic, DCTCP, BBR and TCP Vegas. Gaoxiong Zeng, Wei Bai 0001, Kai Chen 0005, Dongsu Han, Yibo Zhu 0001 |
IEEE/ACM Trans. Netw. | 6 |
| 2022 | DeepCC: Bridging the Gap Between Congestion Control and Applications via Multiobjective OptimizationabstractThe increasingly complicated and diverse applications have distinct network performance demands, e.g., some desire high throughput while others require low latency. Traditional congestion controls (CC) have no perception of these demands. Consequently, literatures have explored the objective-specific algorithms, which are based on either offline training or online learning, to adapt to certain application demands. However, once generated, such algorithms are tailored to a specific performance objective function. Newly emerged performance demands in a changeable network environment require either expensive retraining (in the case of offline training), or manually redesigning a new objective function (in the case of online learning). To address this problem, we propose a novel architecture, DeepCC. It generates a CC agent that is generically applicable to a wide range of application requirements and network conditions. The key idea of DeepCC is to leverage both offline deep reinforcement learning and online fine-tuning. In the offline phase, instead of training towards a specific objective function, DeepCC trains its deep neural network model using multi-objective optimization. With the trained model, DeepCC offers near Pareto optimal policies w.r.t different user-specified trade-offs between throughput, delay, and loss rate without any redesigning or retraining. In addition, a quick online fine-tuning phase further helps DeepCC achieve the application-specific demands under dynamic network conditions. The simulation and real-world experiments show that DeepCC outperforms state-of-the-art schemes in a wide range of settings. DeepCC gains a higher target completion ratio of application requirements up to 67.4% than that of other schemes, even in an untrained environment. Lei Zhang 0157, Yong Cui 0001, Mowei Wang, Kewei Zhu, Yibo Zhu 0001, Yong Jiang 0001 |
IEEE/ACM Trans. Netw. | 5 |
| 2021 | Towards timeout-less transport in commodity datacenter networksabstractDespite recent advances in datacenter networks, timeouts caused by congestion packet losses still remain a major cause of high tail latency. Priority-based Flow Control (PFC) was introduced to make the network lossless, but its Head-of-Line blocking nature causes various performance and management problems. In this paper, we ask if it is possible to design a network that achieves (near) zero timeout only using commodity hardware in datacenters. Hwijoon Lim, Wei Bai 0001, Yibo Zhu 0001, Youngmok Jung, Dongsu Han |
EuroSys | 3 |
| 2021 | AutoLRS: Automatic Learning-Rate Schedule by Bayesian Optimization on the Fly
Tianyi Zhou 0001, Liangyu Zhao, Yibo Zhu 0001, Chuanxiong Guo, Marco Canini, Arvind Krishnamurthy |
ICLR | 4 |
| 2020 | Elastic parameter server load distribution in deep learning clustersabstractIn distributed DNN training, parameter servers (PS) can become performance bottlenecks due to PS stragglers, caused by imbalanced parameter distribution, bandwidth contention, or computation interference. Few existing studies have investigated efficient parameter (aka load) distribution among PSs. We observe significant training inefficiency with the current parameter assignment in representative machine learning frameworks (e.g., MXNet, TensorFlow), and big potential for training acceleration with better PS load distribution. We design PSLD, a dynamic parameter server load distribution scheme, to mitigate PS straggler issues and accelerate distributed model training in the PS architecture. An exploitation-exploration method is carefully designed to scale in and out parameter servers and adjust parameter distribution among PSs on the go. We also design an elastic PS scaling module to carry out our scheme with little interruption to the training process. We implement our module on top of open-source PS architectures, including MXNet and BytePS. Testbed experiments show up to 2.86x speed-up in model training with PSLD, for different ML models under various straggler settings. Yangrui Chen, Yanghua Peng, Yixin Bao, Chuan Wu 0001, Yibo Zhu 0001, Chuanxiong Guo |
SoCC | 5 |
| 2020 | PipeSwitch: Fast Pipelined Context Switching for Deep Learning Applications
Zhihao Bai, Zhen Zhang 0063, Yibo Zhu 0001, Xin Jin 0008 |
OSDI | 3 |
| 2020 | A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU Clusters
Yibo Zhu 0001, Chang Lan, Bairen Yi, Yong Cui 0001, Chuanxiong Guo |
OSDI | 2 |
| 2020 | TEA: Enabling State-Intensive Network Functions on Programmable SwitchesabstractProgrammable switches have been touted as an attractive alternative for deploying network functions (NFs) such as network address translators (NATs), load balancers, and firewalls. However, their limited memory capacity has been a major stumbling block that has stymied their adoption for supporting state-intensive NFs such as cloud-scale NATs and load balancers that maintain millions of flow-table entries. In this paper, we explore a new approach that leverages DRAM on servers available in typical NFV clusters. Our new system architecture, called TEA (Table Extension Architecture), provides a virtual table abstraction that allows NFs on programmable switches to look up large virtual tables built on external DRAM. Our approach enables switch ASICs to access external DRAM purely in the data plane without involving CPUs on servers. We address key design and implementation challenges in realizing this idea. We demonstrate its feasibility and practicality with our implementation on a Tofino-based programmable switch. Our evaluation shows that NFs built with TEA can look up table entries on external DRAM with low and predictable latency (1.8-2.2 μs) and the lookup throughput can be linearly scaled with additional servers (138 million lookups per seconds with 8 servers). Daehyeok Kim, Zaoxing Liu, Yibo Zhu 0001, Changhoon Kim, Jeongkeun Lee, Vyas Sekar, Srinivasan Seshan |
SIGCOMM | 3 |
| 2019 | Congestion Control for Cross-Datacenter NetworksabstractGeographically distributed applications hosted on cloud are becoming prevalent. They run on cross-datacenter network that consists of multiple data center networks (DCNs) connected by a wide area network (WAN). Such a cross-DC network imposes significant challenges in transport design because the DCN and WAN segments have vastly distinct characteristics (e.g., butter depths, RTTs). In this paper, we find that existing DCN or WAN transports reacting to ECN or delay alone do not (and cannot be extended to) work well for such an environment. The key reason is that neither of the signals, by itself, can simultaneously capture the location and degree of congestion. This is due to the discrepancies between DCN and WAN. Motivated by this, we present the design and implementation of GEMINI that strategically integrates both ECN and delay signals for cross-DC congestion control. To achieve low latency, GEMINI bounds the inter-DC latency with delay signal and prevents the intra-DC packet loss with ECN. To maintain high throughput, GEMINI modulates the window dynamics and maintains low butter occupancy utilizing both congestion signals. GEMINI is implemented in Linux kernel and evaluated by extensive testbed experiments. Results show that GEMINI achieves up to 53%, 31% and 76% reduction of small flow average completion times compared to TCP Cubic, DCTCP and BBR; and up to 58% reduction of large flow average completion times compared to TCP Vegas. Gaoxiong Zeng, Wei Bai 0001, Kai Chen 0005, Dongsu Han, Yibo Zhu 0001 |
ICNP | 6 |
| 2019 | Tiresias: A GPU Cluster Manager for Distributed Deep Learning
Juncheng Gu, Mosharaf Chowdhury, Kang G. Shin, Yibo Zhu 0001, Myeongjae Jeon, Junjie Qian, Hongqiang Harry Liu, Chuanxiong Guo |
NSDI | 4 |
| 2019 | FreeFlow: Software-based Virtual RDMA Networking for Containerized Clouds
Daehyeok Kim, Tianlong Yu, Hongqiang Harry Liu, Yibo Zhu 0001, Jitendra Padhye, Shachar Raindel, Chuanxiong Guo, Vyas Sekar, Srinivasan Seshan |
NSDI | 4 |
| 2019 | Slim: OS Kernel Support for a Low-Overhead Container Overlay Network
Danyang Zhuo, Kaiyuan Zhang 0001, Yibo Zhu 0001, Hongqiang Harry Liu, Matthew Rockett, Arvind Krishnamurthy, Thomas E. Anderson |
NSDI | 3 |
| 2019 | A generic communication scheduler for distributed DNN training accelerationabstractWe present ByteScheduler, a generic communication scheduler for distributed DNN training acceleration. ByteScheduler is based on our principled analysis that partitioning and rearranging the tensor transmissions can result in optimal results in theory and good performance in real-world even with scheduling overhead. To make ByteScheduler work generally for various DNN training frameworks, we introduce a unified abstraction and a Dependency Proxy mechanism to enable communication scheduling without breaking the original dependencies in framework engines. We further introduce a Bayesian Optimization approach to auto-tune tensor partition size and other parameters for different training models under various networking conditions. ByteScheduler now supports TensorFlow, PyTorch, and MXNet without modifying their source code, and works well with both Parameter Server (PS) and all-reduce architectures for gradient synchronization, using either TCP or RDMA. Our experiments show that ByteScheduler accelerates training with all experimented system configurations and DNN models, by up to 196% (or 2.96X of original speed). Yanghua Peng, Yibo Zhu 0001, Yangrui Chen, Yixin Bao, Bairen Yi, Chang Lan, Chuan Wu 0001, Chuanxiong Guo |
SOSP | 2 |
| 2019 | Tagger: Practical PFC Deadlock Prevention in Data Center NetworksabstractRemote direct memory access over converged Ethernet deployments is vulnerable to deadlocks induced by priority flow control. Prior solutions for deadlock prevention either require significant changes to routing protocols or require excessive buffers in the switches. In this paper, we propose Tagger, a scheme for deadlock prevention. It does not require any changes to the routing protocol and needs only modest buffers. Tagger is based on the insight that given a set of expected lossless routes, a simple tagging scheme can be developed to ensure that no deadlock will occur under any failure conditions. Packets that do not travel on these lossless routes may be dropped under extreme conditions. We design such a scheme, prove that it prevents deadlock, and implement it efficiently on commodity hardware. Shuihai Hu, Yibo Zhu 0001, Peng Cheng 0005, Chuanxiong Guo, Jitendra Padhye, Kai Chen 0005 |
IEEE/ACM Trans. Netw. | 2 |
| 2018 | Generic External Memory for Switch Data PlanesabstractNetwork switches are an attractive vantage point to serve various network applications and functions such as load balancing and virtual switching because of their in-network location and high packet processing rate. Recent advances in programmable switch ASICs open more opportunities for offloading various functionality to switches. However, the limited memory capacity on switches has been a major challenge that such applications struggle to deal with. In this paper, we envision that by enabling network switches to access remote memory purely from data planes, the performance of a wide range of applications can be improved. We design three remote memory primitives, leveraging RDMA operations, and show the feasibility of accessing remote memory from switches using our prototype implementation. Daehyeok Kim, Yibo Zhu 0001, Changhoon Kim, Jeongkeun Lee, Srinivasan Seshan |
HotNets | 2 |
| 2018 | 007: Democratically Finding the Cause of Packet Drops
Behnaz Arzani, Selim Ciraci, Luiz F. O. Chamon, Yibo Zhu 0001, Hongqiang Liu, Jitendra Padhye, Boon Thau Loo, Geoff Outhred |
NSDI | 4 |
| 2018 | Hyperloop: group-based NIC-offloading to accelerate replicated transactions in multi-tenant storage systemsabstractStorage systems in data centers are an important component of large-scale online services. They typically perform replicated transactional operations for high data availability and integrity. Today, however, such operations suffer from high tail latency even with recent kernel bypass and storage optimizations, and thus affect the predictability of end-to-end performance of these services. We observe that the root cause of the problem is the involvement of the CPU, a precious commodity in multi-tenant settings, in the critical path of replicated transactions. In this paper, we present HyperLoop, a new framework that removes CPU from the critical path of replicated transactions in storage systems by offloading them to commodity RDMA NICs, with non-volatile memory as the storage medium. To achieve this, we develop new and general NIC offloading primitives that can perform memory operations on all nodes in a replication group while guaranteeing ACID properties without CPU involvement. We demonstrate that popular storage applications can be easily optimized using our primitives. Our evaluation results with microbenchmarks and application benchmarks show that HyperLoop can reduce 99th percentile latency ≈ 800X with close to 0% CPU consumption on replicas. Daehyeok Kim, Amir Saman Memaripour, Anirudh Badam, Yibo Zhu 0001, Hongqiang Harry Liu, Jitendra Padhye, Shachar Raindel, Steven Swanson, Vyas Sekar, Srinivasan Seshan |
SIGCOMM | 4 |
| 2017 | Combining ECN and RTT for Datacenter TransportabstractDatacenter transports should provide low average and tail flow completion times (FCT) to achieve desired application performance. While most prior datacenter transports take either ECN or RTT as congestion signal, this paper makes a case that both signals are indispensable: ECN, as a per-hop signal, is more effective to prevent packet loss; while RTT, as an end-to-end signal, controls end-to-end queueing delay better. As persistent low flow completion times imply low queueing delay and near zero packet loss, we introduce EAR, a new datacenter transport that hears and reacts to both ECN and RTT. Our preliminary results show that: 1) compared to delay-based DCTCP, EAR achieves up to 91% lower packet losses and 93% fewer timeouts; 2) compared to ECN-based DCTCP, EAR reduces RTT by up to 32% for cross-rack traffic in a 4-level fattree. As a result, EAR delivers persistent low average and tail completion times under various scenarios in large scale simulations. Gaoxiong Zeng, Wei Bai 0001, Kai Chen 0005, Dongsu Han, Yibo Zhu 0001 |
APNet | 6 |
| 2017 | CrystalNet: Faithfully Emulating Large Production NetworksabstractNetwork reliability is critical for large clouds and online service providers like Microsoft. Our network is large, heterogeneous, complex and undergoes constant churns. In such an environment even small issues triggered by device failures, buggy device software, configuration errors, unproven management tools and unavoidable human errors can quickly cause large outages. A promising way to minimize such network outages is to proactively validate all network operations in a high-fidelity network emulator, before they are carried out in production. To this end, we present CrystalNet, a cloud-scale, high-fidelity network emulator. It runs real network device firmwares in a network of containers and virtual machines, loaded with production configurations. Network engineers can use the same management tools and methods to interact with the emulated network as they do with a production network. CrystalNet can handle heterogeneous device firmwares and can scale to emulate thousands of network devices in a matter of minutes. To reduce resource consumption, it carefully selects a boundary of emulations, while ensuring correctness of propagation of network changes. Microsoft's network engineers use CrystalNet on a daily basis to test planned network operations. Our experience shows that CrystalNet enables operators to detect many issues that could trigger significant outages. Hongqiang Harry Liu, Yibo Zhu 0001, Jitendra Padhye, Jiaxin Cao, Sri Tallapragada, Nuno P. Lopes, Andrey Rybalchenko, Guohan Lu |
SOSP | 2 |
| 2016 | ECN or Delay: Lessons Learnt from Analysis of DCQCN and TIMELYabstractData center networks, and especially drop-free RoCEv2 networks require efficient congestion control protocols. DCQCN (ECN-based) and TIMELY (delay-based) are two recent proposals for this purpose. In this paper, we analyze DCQCN and TIMELY using fluid models and simulations, for stability, convergence, fairness and flow completion time. We uncover several surprising behaviors of these protocols. For example, we show that DCQCN exhibits non-monotonic stability behavior, and that TIMELY can converge to stable regime with arbitrary unfairness. We propose simple fixes and tuning for ensuring that both protocols converge to and are stable at the fair share point. Finally, using lessons learnt from the analysis, we address the broader question: are there fundamental reasons to prefer either ECN or delay for end-to-end congestion control in data center networks? We argue that ECN is a better congestion signal, due to the way modern switches mark packets, and due to a fundamental limitation of end-to-end delay-based protocols, that we derive. Yibo Zhu 0001, Manya Ghobadi, Vishal Misra, Jitendra Padhye |
CoNEXT | 1 |
| 2016 | Trimming the Smartphone Network StackabstractNetwork transmissions are the cornerstone of most mobile apps today, and a main contributor to energy consumption. We use a componentized energy model to quantify energy use by device, and observe significant energy consumption by the CPU in network operations. We assert that optimizing network operations in the CPU can produce significant energy savings, and explore the impact of two potential approaches: one-copy data moves and offloading the network stack to the basestation. Yanzi Zhu, Yibo Zhu 0001, Ana Nika, Ben Y. Zhao, Haitao Zheng 0001 |
HotNets | 2 |
| 2016 | Empirical Validation of Commodity Spectrum MonitoringabstractWe describe our efforts to empirically validate a distributed spectrum monitoring system built on commodity smartphones and embedded low-cost spectrum sensors. This system enables real-time spectrum sensing, identifies and locates active transmitters, and generates alarm events when detecting anomalous transmitters. To evaluate the feasibility of such a platform, we perform detailed experiments using a prototype hardware platform using smartphones and RTL dongles. We identify multiple sources of error in the sensing results and the end-user overhead (i.e. smartphone energy draw). We propose and implement a variety of techniques to identify and overcome errors and uncertainty in the data, and to reduce energy consumption. Our work demonstrates the basic viability of user-driven spectrum monitoring on commodity devices. Ana Nika, Zhijing Li 0001, Yanzi Zhu, Yibo Zhu 0001, Ben Y. Zhao, Haitao Zheng 0001 |
SenSys | 4 |
| 2015 | Reusing 60GHz Radios for Mobile Radar ImagingabstractThe future of mobile computing involves autonomous drones, robots and vehicles. To accurately sense their surroundings in a variety of scenarios, these mobile computers require a robust environmental mapping system. One attractive approach is to reuse millimeterwave communication hardware in these devices, e.g. 60GHz networking chipset, and capture signals reflected by the target surface. The devices can also move while collecting reflection signals, creating a large synthetic aperture radar (SAR) for high-precision RF imaging. Our experimental measurements, however, show that this approach provides poor precision in practice, as imaging results are highly sensitive to device positioning errors that translate into phase errors. We address this challenge by proposing a new 60GHz imaging algorithm, {\em RSS Series Analysis}, which images an object using only RSS measurements recorded along the device's trajectory. In addition to object location, our algorithm can discover a rich set of object surface properties at high precision, including object surface orientation, curvature, boundaries, and surface material. We tested our system on a variety of common household objects (between 5cm--30cm in width). Results show that it achieves high accuracy (cm level) in a variety of dimensions, and is highly robust against noises in device position and trajectory tracking. We believe that this is the first practical mobile imaging system (re)using 60GHz networking devices, and provides a basic primitive towards the construction of detailed environmental mapping systems. Yanzi Zhu, Yibo Zhu 0001, Ben Y. Zhao, Haitao Zheng 0001 |
MobiCom | 2 |
| 2015 | Congestion Control for Large-Scale RDMA DeploymentsabstractModern datacenter applications demand high throughput (40Gbps) and ultra-low latency (< 10 μs per hop) from the network, with low CPU overhead. Standard TCP/IP stacks cannot meet these requirements, but Remote Direct Memory Access (RDMA) can. On IP-routed datacenter networks, RDMA is deployed using RoCEv2 protocol, which relies on Priority-based Flow Control (PFC) to enable a drop-free network. However, PFC can lead to poor application performance due to problems like head-of-line blocking and unfairness. To alleviates these problems, we introduce DCQCN, an end-to-end congestion control scheme for RoCEv2. To optimize DCQCN performance, we build a fluid model, and provide guidelines for tuning switch buffer thresholds, and other protocol parameters. Using a 3-tier Clos network testbed, we show that DCQCN dramatically improves throughput and fairness of RoCEv2 RDMA traffic. DCQCN is implemented in Mellanox NICs, and is being deployed in Microsoft's datacenters. Yibo Zhu 0001, Haggai Eran, Daniel Firestone, Chuanxiong Guo, Marina Lipshteyn, Yehonatan Liron, Jitendra Padhye, Shachar Raindel, Mohamad Haj Yahia, Ming Zhang 0005 |
SIGCOMM | 1 |
| 2015 | Packet-Level Telemetry in Large Datacenter NetworksabstractDebugging faults in complex networks often requires capturing and analyzing traffic at the packet level. In this task, datacenter networks (DCNs) present unique challenges with their scale, traffic volume, and diversity of faults. To troubleshoot faults in a timely manner, DCN administrators must a) identify affected packets inside large volume of traffic; b) track them across multiple network components; c) analyze traffic traces for fault patterns; and d) test or confirm potential causes. To our knowledge, no tool today can achieve both the specificity and scale required for this task. Yibo Zhu 0001, Nanxi Kang, Jiaxin Cao, Albert G. Greenberg, Guohan Lu, Ratul Mahajan, David A. Maltz, Ming Zhang 0005, Ben Y. Zhao, Haitao Zheng 0001 |
SIGCOMM | 1 |
| 2015 | Energy and Performance of Smartphone Radio Bundling in Outdoor EnvironmentsabstractMost of today's mobile devices come equipped with both cellular LTE and WiFi wireless radios, making radio bundling (simultaneous data transfers over multiple interfaces) both appealing and practical. Despite recent studies documenting the benefits of radio bundling with MPTCP, many fundamental questions remain about potential gains from radio bundling, or the relationship between performance and energy consumption in these scenarios. In this study, we seek to answer these questions using extensive measurements to empirically characterize both energy and performance for radio bundling approaches. In doing so, we quantify potential gains of bundling using MPTCP versus an ideal protocol. We study the links between traffic partitioning and bundling performance, and use a novel componentized energy model to quantify the energy consumed by CPUs (and radios) during traffic management. Our results show that MPTCP achieves only a fraction of the total performance gain possible, and that its energy-agnostic design leads to considerable power consumption by the CPU. We conclude that not only there is room for improved bundling performance, but an energy-aware bundling protocol is likely to achieve a much better tradeoff between performance and power consumption. Ana Nika, Yibo Zhu 0001, Ning Ding 0004, Abhilash Jindal, Y. Charlie Hu, Ben Y. Zhao, Haitao Zheng 0001 |
WWW | 2 |
| 2014 | Demystifying 60GHz outdoor picocellsabstractMobile network traffic is set to explode in our near future, driven by the growth of bandwidth-hungry media applications. Current capacity solutions, including buying spectrum, WiFi offloading, and LTE picocells, are unlikely to supply the orders-of-magnitude bandwidth increase we need. In this paper, we explore a dramatically different alternative in the form of 60GHz mmwave picocells with highly directional links. While industry is investigating other mmwave bands (e.g. 28GHz to avoid oxygen absorption), we prefer the unlicensed 60GHz band with highly directional, short-range links (~100m). 60GHz links truly reap the spatial reuse benefits of small cells while delivering high per-user data rates and leveraging efforts on indoor 60GHz PHY technology and standards. Using extensive measurements on off-the-shelf 60GHz radios and system-level simulations, we explore the feasibility of 60GHz picocells by characterizing range, attenuation due to reflections, sensitivity to movement and blockage, and interference in typical urban environments. Our results dispel some common myths, and show that there are no fundamental physical barriers to high-capacity 60GHz outdoor picocells. We conclude by identifying open challenges and associated research opportunities. Yibo Zhu 0001, Zengbin Zhang, Zhinus Marzi, Chris Nelson, Upamanyu Madhow, Ben Y. Zhao, Haitao Zheng 0001 |
MobiCom | 1 |
| 2014 | Cutting the cord: a robust wireless facilities network for data centersabstractToday's network control and management traffic are limited by their reliance on existing data networks. Fate sharing in this context is highly undesirable, since control traffic has very different availability and traffic delivery requirements. In this paper, we explore the feasibility of building a dedicated wireless facilities network for data centers. We propose Angora, a low-latency facilities network using low-cost, 60GHz beamforming radios that provides robust paths decoupled from the wired network, and flexibility to adapt to workloads and network dynamics. We describe our solutions to address challenges in link coordination, link interference and network failures. Our testbed measurements and simulation results show that Angora enables large number of low-latency control paths to run concurrently, while providing low latency end-to-end message delivery with high tolerance for radio and rack failures. Yibo Zhu 0001, Zengbin Zhang, Amin Vahdat, Ben Y. Zhao, Haitao Zheng 0001 |
MobiCom | 1 |
| 2013 | Datacast: A Scalable and Efficient Reliable Group Data Delivery Service for Data CentersabstractReliable Group Data Delivery (RGDD) is a pervasive traffic pattern in data centers. In an RGDD group, a sender needs to reliably deliver a copy of data to all the receivers. Existing solutions either do not scale due to the large number of RGDD groups (e.g., IP multicast) or cannot efficiently use network bandwidth (e.g., end-host overlays). Motivated by recent advances on data center network topology designs (multiple edge-disjoint Steiner trees for RGDD) and innovations on network devices (practical in-network packet caching), we propose Datacast for RGDD. Datacast explores two design spaces: 1) Datacast uses multiple edge-disjoint Steiner trees for data delivery acceleration. 2) Datacast leverages in-network packet caching and introduces a simple soft-state based congestion control algorithm to address the scalability and efficiency issues of RGDD. Our analysis reveals that Datacast congestion control works well with small cache sizes (e.g., 125KB) and causes few duplicate data transmissions (e.g., 1.19%). Both simulations and experiments confirm our theoretical analysis. We also use experiments to compare the performance of Datacast and BitTorrent. In a BCube(4, 1) with 1Gbps links, we use both Datacast and BitTorrent to transmit 4GB data. The link stress of Datacast is 1.01, while it is 1.39 for BitTorrent. By using two Steiner trees, Datacast finishes the transmission in 16.9s, while BitTorrent uses 52s. Jiaxin Cao, Chuanxiong Guo, Guohan Lu, Yongqiang Xiong, Yixin Zheng, Yongguang Zhang, Yibo Zhu 0001, Chen Chen 0019, Ye Tian 0004 |
IEEE J. Sel. Areas Commun. | 7 |
| 2012 | Datacast: a scalable and efficient reliable group data delivery service for data centersabstractReliable Group Data Delivery (RGDD) is a pervasive traffic pattern in data centers. In an RGDD group, a sender needs to reliably deliver a copy of data to all the receivers. Existing solutions either do not scale due to the large number of RGDD groups (e.g., IP multicast) or cannot efficiently use network bandwidth (e.g., end-host overlays). Jiaxin Cao, Chuanxiong Guo, Guohan Lu, Yongqiang Xiong, Yixin Zheng, Yongguang Zhang, Yibo Zhu 0001, Chen Chen 0019 |
CoNEXT | 7 |
| 2012 | Mirror mirror on the ceiling: flexible wireless links for data centersabstractModern data centers are massive, and support a range of distributed applications across potentially hundreds of server racks. As their utilization and bandwidth needs continue to grow, traditional methods of augmenting bandwidth have proven complex and costly in time and resources. Recent measurements show that data center traffic is often limited by congestion loss caused by short traffic bursts. Thus an attractive alternative to adding physical bandwidth is to augment wired links with wireless links in the 60 GHz band. Zengbin Zhang, Yibo Zhu 0001, Saipriya Kumar, Amin Vahdat, Ben Y. Zhao, Haitao Zheng 0001 |
SIGCOMM | 3 |
| 2012 | Serf and turf: crowdturfing for fun and profitabstractPopular Internet services in recent years have shown that remarkable things can be achieved by harnessing the power of the masses using crowd-sourcing systems. However, crowd-sourcing systems can also pose a real challenge to existing security mechanisms deployed to protect Internet services. Many of these security techniques rely on the assumption that malicious activity is generated automatically by automated programs. Thus they would perform poorly or be easily bypassed when attacks are generated by real users working in a crowd-sourcing system. Through measurements, we have found surprising evidence showing that not only do malicious crowd-sourcing systems exist, but they are rapidly growing in both user base and total revenue. We describe in this paper a significant effort to study and understand these "crowdturfing" systems in today's Internet. We use detailed crawls to extract data about the size and operational structure of these crowdturfing systems. We analyze details of campaigns offered and performed in these sites, and evaluate their end-to-end effectiveness by running active, benign campaigns of our own. Finally, we study and compare the source of workers on crowdturfing sites in different countries. Our results suggest that campaigns on these systems are highly effective at reaching users, and their continuing growth poses a concrete threat to online communities both in the US and elsewhere. Gang Wang 0011, Christo Wilson, Xiaohan Zhao, Yibo Zhu 0001, Manish Mohanlal, Haitao Zheng 0001, Ben Y. Zhao |
WWW | 4 |
| 2011 | Tarantula: Towards an Accurate Network Coordinate System by Handling Major Portion of TIVsabstractNetwork Coordinate (NC) systems provide an efficient and scalable mechanism to estimate latencies among hosts. However, many popular algorithms like Vivaldi suffer greatly from the existence of Triangle Inequality Violations (TIVs). Two-layer systems like Pharos and hierarchical Vivaldi have been proposed to remedy the impact of TIVs. They divide the whole space into several location-based clusters and run NC systems on both global layer and local layer. However, the two-layer model is only able to optimize the intra-cluster links relating to a limited portion of TIV triangles. In this paper, we propose a new NC system, Tarantula, which divides the space in a novel way. By categorizing the TIVs into three classes, we show that Tarantula handles a much larger portion of existing TIVs than two-layer systems. Moreover, we present two techniques to further strengthen the Tarantula system: 1) relate the updating step size in the Vivaldi algorithm used in Tarantula to ground-truth latency so as to improve the prediction for short links; 2) propose Dynamic Cluster Optimization to dynamically adjust clustering of hosts. Our experimental results show that Tarantula outperforms Pharos and Vivaldi significantly in terms of estimation accuracy. When implementing different NC systems in the application of server selection and detour finding, Tarantula again performs the best. Yang Chen 0001, Yibo Zhu 0001, Cong Ding 0001, Beixing Deng, Xing Li 0001 |
GLOBECOM | 3 |