EDBT 2026 Demo / reviewers in the wild / expert
Peng Cheng 0005
dblp:76/185-5
· DBLP profile ↗
55ranked-venue papers
1as first author
34since 2021 · last 2026
0000-0003-4014-4757ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 32 · 1 first-author · 17 since 2021Systems, architecture and hardware · 15 · 9 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-trainingabstractKailai Yang, Xiao Liu, Lei Ji, Hao Li, Xiao Liang, Zhiwei Liu, Yeyun Gong, Peng Cheng, Mao Yang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Kailai Yang, Xiao Liu 0029, Lei Ji 0001, Hao Li 0074, Zhiwei Liu 0003, Yeyun Gong, Peng Cheng 0005, Mao Yang 0004 |
ACL (1) | 8 |
| 2026 | OptiFlow: Towards LLM-Driven Optimization of Collective Communication Algorithms
Ziyue Yang 0002, Kaihui Gao, Shuai Wang 0028, Li Chen 0008, Zhixiong Niu, Ran Shu 0001, Wenxue Cheng, Peng Cheng 0005, Yongqiang Xiong, Dan Li 0001 |
APNet | 9 |
| 2026 | MSCCL++: Rethinking GPU Communication Abstractions for AI InferenceabstractAI applications increasingly run on fast-evolving, heterogeneous hardware to maximize performance, but general-purpose libraries lag in supporting these features. Performance-minded programmers often build custom communication stacks that are fast but error-prone and non-portable. This paper introduces MSCCL++, a design methodology for developing high-performance, portable communication kernels. It provides (1) a low-level, performance-preserving primitive interface that exposes minimal hardware abstractions while hiding the complexities of synchronization and consistency, (2) a higher-level DSL for application developers to implement workload-specific communication algorithms, and (3) a library of efficient algorithms implementing the standard collective API, enabling adoption by users with minimal expertise. Compared to state-of-the-art baselines, MSCCL++ achieves geomean speedups of 1.7× (up to 5.4×) for collective communication and 1.2× (up to 1.38×) for AI inference workloads. MSCCL++ is in production of multiple AI services provided by Microsoft Azure, and has also been adopted by RCCL, the GPU collective communication library maintained by AMD. MSCCL++ is open source and available at https://github.com/microsoft/mscclpp. Our two years of experience with MSCCL++ suggests that its abstractions are robust, enabling support for new hardware features, such as multimem, within weeks of development. Changho Hwang, Peng Cheng 0005, Roshan Dathathri, Abhinav Jangda, Saeed Maleki, Madan Musuvathi, Olli Saarikivi, Aashaka Shah, Ziyue Yang 0002, Binyang Li, Caio Rocha, Mahdieh Ghazimirsaeed, Sreevatsa Anantharamu |
ASPLOS (2) | 2 |
| 2026 | SmartNIC-Enabled Live Migration for Storage-Optimized VMs with PYROCUMULUS
Jiechen Zhao 0002, Ran Shu 0001, Ziyue Yang 0002, Rui Ma 0021, Derek Chiou, Natalie D. Enright Jerger, Peng Cheng 0005, Yongqiang Xiong |
NSDI | 8 |
| 2026 | SuperBench: A Proactive Validation System for Improving Reliability of Cloud AI InfrastructureabstractReliability in cloud AI infrastructure is crucial for cloud service providers, prompting the widespread use of hardware redundancies. However, these redundancies can inadvertently lead to hidden degradation, known as “gray failure”, for AI workloads, significantly affecting end-to-end performance and concealing performance issues, which complicates root cause analysis for failures and regressions. We introduce SuperBench, a proactive validation system for AI infrastructure that mitigates hidden degradation caused by hardware redundancies and enhances overall reliability. SuperBench features a comprehensive benchmark suite, capable of evaluating individual hardware components and representing most real AI workloads. It comprises a Validator that learns benchmark criteria to pinpoint defective components clearly. Additionally, SuperBench incorporates a Selector to balance validation time and issue-related penalties, enabling optimal timing for validation execution with a tailored subset of benchmarks. Through testbed evaluation and simulation, we demonstrate that SuperBench can increase the mean time between incidents by up to 22.61×. SuperBench has been successfully deployed in Azure production, validating hundreds of thousands of GPUs every year. Yifan Xiong 0001, Ziyue Yang 0002, Guoshuai Zhao 0001, Dong Zhong, Boris Pinzur, Jie Zhang 0048, Yang Wang 0053, Hossein Pourreza, Jeff Baxter, Kushal Datta, Prabhat Ram, Luke Melton, Joe Chau, Peng Cheng 0005, Yongqiang Xiong, Lidong Zhou |
ACM Trans. Comput. Syst. | 18 |
| 2026 | Fine-Grained Scheduling of In-Network Aggregation Resources for Efficient Machine Learning Service
Shichen Dong, Zhixiong Niu, Mingchao Zhang, Zhiying Xu, Chuntao Hu, Pengzhi Zhu, Qingchun Song, Peng Cheng 0005, Cam-Tu Nguyen, Shaoling Sun, Xiaohu Xu, Yongqiang Xiong, Wei Wang 0002, Xiaoliang Wang 0001, Guihai Chen |
IEEE Trans. Netw. | 9 |
| 2025 | RapidScribe: Bandwidth-aware Parallel Checkpoint for Distributed Neural-Network TrainingabstractAs machine learning models grow in complexity and size, necessitating the use of extensive GPU clusters, the challenge of managing frequent system failures becomes increasingly critical. These failures can lead to substantial losses in computational time and resources. Large-scale model training often does not use GPU redundancy for fault tolerance, underscoring the necessity for an efficient checkpoint mechanism for GPU failure recovery. Enhanced failure resilience demands more frequent checkpoints; however suspend-and-resume based checkpoints could severely lower training throughput.RapidScribe addresses this challenge head-on by leveraging unused GPU bus bandwidth to facilitate tensor transfers from GPU to host in concurrent with ongoing training. At the heart of RapidScribe is the Steno protocol, which lowers checkpoint traffic aggressively by copying both gradients and model training states. Steno generates copy schedules to be fully interleaved within training processes to minimize training interruptions. The low bandwidth usage by Steno makes it even a viable solution for disk-based checkpoints. Experiments show that RapidScribe’s parallel checkpoint techniques ensure that frequent checkpoint operations are deeply synchronized with neural network training cycles, preserving near-baseline throughput across various distributed training configurations on a AMD MI250 GPU cluster, thus providing a robust checkpoint-based solution to GPU fault tolerance. Shuotao Xu, Yuqing Yang 0001, Peng Cheng 0005 |
ICDCS | 6 |
| 2025 | Automated Proof Generation for Rust Code via Self-EvolutionabstractEnsuring correctness is crucial for code generation. Formal verification offers a
definitive assurance of correctness, but demands substantial human effort in proof
construction and hence raises a pressing need for automation. The primary obsta-
cle lies in the severe lack of data—there is much fewer proofs than code snippets
for Large Language Models (LLMs) to train upon. In this paper, we introduce
SAFE, a framework that overcomes the lack of human-written proofs to enable
automated proof generation of Rust code. SAFE establishes a self-evolving cycle
where data synthesis and fine-tuning collaborate to enhance the model capability,
leveraging the definitive power of a symbolic verifier in telling correct proofs from
incorrect ones. SAFE also re-purposes the large number of synthesized incorrect
proofs to train the self-debugging capability of the fine-tuned models, empowering
them to fix incorrect proofs based on the verifier’s feedback. SAFE demonstrates
superior efficiency and precision compared to GPT-4o. Through tens of thousands
of synthesized proofs and the self-debugging mechanism, we improve the capa-
bility of open-source models, initially unacquainted with formal verification, to
automatically write proofs for Rust code. This advancement leads to a signifi-
cant improvement in performance, achieving a 52.52% accuracy rate in a bench-
mark crafted by human experts, a significant leap over GPT-4o’s performance of
14.39%. Tianyu Chen 0006, Shan Lu 0001, Yeyun Gong, Chenyuan Yang, Xuheng Li, Md Rakib Hossain Misu, Hao Yu 0016, Nan Duan 0001, Peng Cheng 0005, Fan Yang 0024, Shuvendu K. Lahiri, Tao Xie 0001, Lidong Zhou |
ICLR | 10 |
| 2025 | Integrative Decoding: Improving Factuality via Implicit Self-consistencyabstractSelf-consistency-based approaches, which involve repeatedly sampling multiple outputs and selecting the most consistent one as the final response, prove to be remarkably effective in improving the factual accuracy of large language models. Nonetheless, existing methods usually have strict constraints on the task format, largely limiting their applicability. In this paper, we present Integrative Decoding (ID), to unlock the potential of self-consistency in open-ended generation tasks. ID operates by constructing a set of inputs, each prepended with a previously sampled response, and then processes them concurrently, with the next token being selected by aggregating of all their corresponding predictions at each decoding step. In essence, this simple approach implicitly incorporates self-consistency in the decoding objective. Extensive evaluation shows that ID consistently enhances factuality over a wide range of language models, with substantial improvements on the TruthfulQA (+11.2%), Biographies (+15.4%) and LongFact (+8.5%) benchmarks. The performance gains amplify progressively as the number of sampled responses increases, indicating the potential of ID to scale up with repeated sampling. Yeyun Gong, Yuji Zhang 0002, Kaishuai Xu, Wenge Liu, Wenjie Li 0002, Jian Jiao 0007, Qi Chen 0009, Peng Cheng 0005, Wayne Xiong |
ICLR | 13 |
| 2025 | Optimizing Large Language Model Training Using FP4 QuantizationabstractThe growing computational demands of training large language models (LLMs) necessitate more efficient methods. Quantized training presents a promising solution by enabling low-bit arithmetic operations to reduce these costs. While FP8 precision has demonstrated feasibility, leveraging FP4 remains a challenge due to significant quantization errors and limited representational capacity. This work introduces the first FP4 training framework for LLMs, addressing these challenges with two key innovations: a differentiable quantization estimator for precise weight updates and an outlier clamping and compensation strategy to prevent activation collapse. To ensure stability, the framework integrates a mixed-precision training scheme and vector-wise quantization. Experimental results demonstrate that our FP4 framework achieves accuracy comparable to BF16 and FP8, with minimal degradation, scaling effectively to 13B-parameter LLMs trained on up to 100B tokens. With the emergence of next-generation hardware supporting FP4, our framework sets a foundation for efficient ultra-low precision training. Yeyun Gong, Xiao Liu 0029, Guoshuai Zhao 0001, Ziyue Yang 0002, Baining Guo, Zhengjun Zha, Peng Cheng 0005 |
ICML | 8 |
| 2025 | Mina: Fine-Grained In-network Aggregation Resource Scheduling for Machine Learning Service
Shichen Dong, Zhixiong Niu, Mingchao Zhang, Zhiying Xu, Chuntao Hu, Pengzhi Zhu, Qingchun Song, Peng Cheng 0005, Cam-Tu Nguyen, Shaoling Sun, Xiaohu Xu, Yongqiang Xiong, Wei Wang 0002, Xiaoliang Wang 0001 |
INFOCOM | 9 |
| 2025 | HyperDrive: Direct Network Telemetry Storage via Programmable SwitchesabstractIn cloud datacenter operations, telemetry and logs are indispensable, enabling essential services such as network diagnostics, auditing, and knowledge discovery. The escalating scale of data centers, coupled with increased bandwidth and finer-grained telemetry, results in an overwhelming volume of data. This proliferation poses significant storage challenges for telemetry systems. In this article, we introduce HyperDrive, an innovative system designed to efficiently store large volumes of telemetry and logs in data centers using programmable switches. This in-network approach effectively mitigates bandwidth bottlenecks commonly associated with traditional endpoint-based methods. To our knowledge, we are the first to use a programmable switch to directly control storage, bypassing the CPU to achieve the best performance. With merely 21% of a switch’s resources, our HyperDrive implementation showcases remarkable scalability and efficiency. Through rigorous evaluation, it has demonstrated linear scaling capabilities, efficiently managing 12 SSDs on a single server with minimal host overhead. In an eight-server testbed, HyperDrive achieved an impressive throughput of approximately 730 Gbps, underscoring its potential to transform data center telemetry and logging practices. Ziyuan Liu 0008, Zhixiong Niu, Ran Shu 0001, Wenxue Cheng, Jacob Nelson 0001, Dan R. K. Ports, Peng Cheng 0005, Yongqiang Xiong |
IEEE Trans. Cloud Comput. | 8 |
| 2025 | Low-Overhead Intra-Host Container Communication With Hardware OffloadingabstractContainers are widely embraced for their deployment and performance benefits over virtual machines. Yet, for many data-intensive applications in containerized clouds, bulky data transfers may impose performance issues. In particular, communication across co-located containers on the same host incurs large overheads in memory copy and the kernel’s TCP stack. Existing solutions such as shared-memory networking and RDMA have their own limitations, including insufficient memory isolation and limited scalability. This paper presents PipeDevice, a new system for low overhead intra-host container communication. PipeDevice follows a hardware-software co-design approach — it offloads data forwarding entirely onto hardware, which accesses application data in hugepages on the host, thereby eliminating CPU overhead from memory copy and TCP processing. PipeDevice preserves memory isolation and scales well to connections, making it deployable in public clouds. Isolation is achieved by allocating dedicated memory to each connection from hugepages. To achieve high scalability, PipeDevice stores the connection states entirely in host DRAM and manages them in software. Evaluation with a prototype implementation on commodity FPGA shows that for delivering 80Gbps across containers PipeDevice saves 63.2% CPU compared to kernel TCP stack, and 40.5% over FreeFlow. PipeDevice provides salient benefits to applications. For example, we port baidu-allreduce to PipeDevice and obtain$\sim 2.2\times $gains in allreduce throughput. Zhixiong Niu, Ran Shu 0001, Peng Cheng 0005, Yongqiang Xiong, Dongsu Han, Chun Jason Xue, Hong Xu 0001 |
IEEE Trans. Netw. | 4 |
| 2024 | NeoMem: Hardware/Software Co-Design for CXL-Native Memory TieringabstractThe Compute Express Link (CXL) interconnect makes it feasible to integrate diverse types of memory into servers via its byte-addressable SerDes links. Considering the various access latency, harnessing the full potential of CXL-based heterogeneous memory systems requires efficient memory tiering. However, prior work can hardly make a fundamental progress owing to low-resolution and high-overhead memory access profiling techniques. To address this critical challenge, we propose a novel memory tiering solution called NeoMem, which features a hardware/software co-design. NeoMem offloads memory profiling functions to CXL device-side controllers, integrating a dedicated hardware unit called NeoProf. NeoProf readily monitors memory accesses and provides the OS with crucial page hotness statistics and other useful system state information. On the OS kernel side, we design a revamped memory-tiering strategy, enabling accurate and timely hot page promotion based on NeoProf statistics. We implement NeoMem on a real FPGA-based CXL memory platform and Linux kernel v6.3. Comprehensive evaluations demonstrate that NeoMem achieves 32% ~ 67% geomean speedup over several existing memory tiering solutions. Zhe Zhou 0002, Tao Zhang 0032, Yang Wang 0053, Ran Shu 0001, Shuotao Xu, Peng Cheng 0005, Yongqiang Xiong, Jie Zhang 0048, Guangyu Sun 0003 |
MICRO | 7 |
| 2024 | SuperBench: Improving Cloud AI Infrastructure Reliability with Proactive Validation
Yifan Xiong 0001, Ziyue Yang 0002, Guoshuai Zhao 0001, Dong Zhong, Boris Pinzur, Jie Zhang 0048, Yang Wang 0053, Hossein Pourreza, Jeff Baxter, Kushal Datta, Prabhat Ram, Luke Melton, Joe Chau, Peng Cheng 0005, Yongqiang Xiong, Lidong Zhou |
USENIX ATC | 18 |
| 2024 | Intelligent Packet Processing for Performant Containers in IoTabstractThis article explores the computing and communication overhead of network processing in Internet of Things (IoT) devices, focusing on containers, a major building block for the edge computing. Our experiments reveal that containers on IoT devices suffer$\sim 2.6\times $higher CPU usage for SoftIRQ processing, ~59% less network throughput, and$2\times $higher per-packet latency on average than native processes. While several existing studies enhance networking performance, they often sacrifice interoperability by requiring special hardware or modifying networking semantics or APIs. Thus, we design and implement a kernel networking accelerator, called SCON, that maintains interoperability, crucial for IoT devices. SCON addresses major bottlenecks in container networking through system-level profiling. We evaluate SCON with three types of IoT devices. On the Raspberry Pi 4, SCON reduces the latencies of major IoT application protocols (e.g., HTTP and MQTT) by$\sim 10\times $, achieving a similar level of latency to the native process. Further analysis shows that SCON reduces CPU usage for SoftIRQ processing by ~26%. We also report similar improvements on the other two IoT devices. Our conclusion is that SCON is unique in significantly reducing the computing and communication overhead of container networking in IoT devices while maintaining interoperability. Furthermore, it works consistently across different types of devices, whether wired or wireless, and regardless of heavy or sporadic traffic. Wonmi Choi, Yeonho Yoo, Kyungwoon Lee, Zhixiong Niu, Peng Cheng 0005, Yongqiang Xiong, Gyeongsik Yang, Chuck Yoo |
IEEE Internet Things J. | 5 |
| 2023 | MINA: Auto-scale In-network Aggregation for Machine Learning Service
Shichen Dong, Zhixiong Niu, Mingchao Zhang, Zhiying Xu, Chuntao Hu, Wei Wang 0002, Pengzhi Zhu, Qingchun Song, Peng Cheng 0005, Yongqiang Xiong, Chen Tian 0001, Cam-Tu Nguyen, Xiaoliang Wang 0001 |
APNet | 10 |
| 2023 | SlimeMold: Hardware Load Balancer at Scale in DatacenterabstractStateful load balancers (LB) are essential services in cloud data centers, playing a crucial role in enhancing the availability and capacity of applications. Numerous studies have proposed methods to improve the throughput, connections per second, and concurrent flows of single LBs. For instance, with the advancement of programmable switches, hardware-based load balancers (HLB) have become mainstream due to their high efficiency. However, programmable switches still face the issue of limited registers and table entries, preventing them from fully meeting the performance requirements of data centers. In this paper, rather than solely focusing on enhancing individual HLBs, we introduce SlimeMold, which enables HLBs to work collaboratively at scale as an integrated LB system in data centers. Ziyuan Liu 0008, Zhixiong Niu, Ran Shu 0001, Guohong Lai, Zongying He, Jacob Nelson 0001, Dan R. K. Ports, Peng Cheng 0005, Yongqiang Xiong |
APNet | 11 |
| 2023 | SegaNet: An Advanced IoT Cloud Gateway for Performant and Priority-Oriented Message DeliveryabstractWith the tremendous growth of IoT, the role of IoT cloud gateways in facilitating communication between IoT devices and the cloud has become more important than ever before. Most previous studies have focused on developing interoperability between IoT and cloud to accommodate various radio protocols. However, they have often neglected the performance aspect of the IoT cloud gateway, leaving users with limited options: either purchasing multiple gateways or connecting only a small number of IoT devices. Through our comprehensive measurements and analysis, we identified five key issues in IoT cloud gateways related to high latency, CPU bottlenecks, inefficient network stacks on ARM, substantial encryption overhead, and the lack of priority support. To address these issues, we propose a new IoT cloud gateway - SegaNet. We carefully design with 1) multiple agents management, 2) efficient TLS encryption, and 3) priority-oriented message delivery. Our prototype evaluation shows up to 16.7 × lower latency and 4.5 × lower CPU consumption than gateways of the existing IoT-cloud ecosystem. Yeonho Yoo, Zhixiong Niu, Chuck Yoo, Peng Cheng 0005, Yongqiang Xiong |
APNet | 4 |
| 2023 | Polaris: Enhancing CXL-based Memory Expanders with Memory-side Prefetching
Zhe Zhou 0002, Shuotao Xu, Tao Zhang 0032, Ran Shu 0001, Peng Cheng 0005, Yongqiang Xiong, Guangyu Sun 0003 |
APPT | 7 |
| 2023 | ElasticFlow: An Elastic Serverless Training Platform for Distributed Deep LearningabstractThis paper proposes ElasticFlow, an elastic serverless training platform for distributed deep learning. ElasticFlow provides a serverless interface with two distinct features: (i) users specify only the deep neural network (DNN) model and hyperparameters for a job, but not the number of GPUs; (ii) users specify the deadline for a job, but not the amount of time to occupy GPUs. In contrast to existing server-centric platforms, ElasticFlow provides performance guarantees in terms of meeting deadlines while alleviating tedious, low-level, and manual resource management for deep learning developers. The characteristics of distributed training introduce two challenges. First, the training throughput scales non-linearly with the number of GPUs. Second, the scaling efficiency is affected by worker placement. To address these challenges, we propose Minimum Satisfactory Share to capture the resource usage of training jobs to meet deadlines, and ElasticFlow performs admission control based on it. We develop a greedy algorithm that dynamically allocates resources to admitted jobs based on diminishing returns. We apply buddy allocation to worker placement to eliminate the effect of topology. Evaluation results on a cluster of 128 GPUs show that ElasticFlow increases the number of jobs that can meet their deadlines by 1.46–7.65× compared to existing solutions. Diandian Gu, Yinmin Zhong, Yifan Xiong 0001, Zhenhua Han, Peng Cheng 0005, Fan Yang 0024, Gang Huang 0001, Xin Jin 0008, Xuanzhe Liu |
ASPLOS (2) | 6 |
| 2023 | Query Processing on Gaming Consolesabstractresearch-article Share on Query Processing on Gaming Consoles Authors: Wei Cui Microsoft Research Asia, CN Microsoft Research Asia, CN 0009-0005-9362-3585View Profile , Qianxi Zhang Microsoft Research Asia, CN Microsoft Research Asia, CN 0000-0002-0646-5365View Profile , Spyros Blanas The Ohio State University, US The Ohio State University, US 0009-0004-2703-7177View Profile , Jesús Camacho-Rodríguez Microsoft, US Microsoft, US 0009-0008-9151-6024View Profile , Brandon Haynes Microsoft Gray Systems Lab, US Microsoft Gray Systems Lab, US 0000-0002-1501-9586View Profile , Yinan Li Microsoft Research, US Microsoft Research, US 0009-0004-5483-2862View Profile , Ravi Ramamurthy Microsoft, USA Microsoft, USA 0000-0002-3484-0038View Profile , Peng Cheng Microsoft Research, CN Microsoft Research, CN 0000-0003-4014-4757View Profile , Rathijit Sen Microsoft, US Microsoft, US 0000-0003-4736-2837View Profile , Matteo Interlandi Microsoft, US Microsoft, US 0000-0002-5756-8321View Profile Authors Info & Claims DaMoN '23: Proceedings of the 19th International Workshop on Data Management on New HardwareJune 2023Pages 86–88https://doi.org/10.1145/3592980.3595313Published:18 June 2023Publication History 0citation191DownloadsMetricsTotal Citations0Total Downloads191Last 12 Months191Last 6 weeks191 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Qianxi Zhang, Spyros Blanas, Jesús Camacho-Rodríguez, Brandon Haynes, Yinan Li 0009, Ravishankar Ramamurthy, Peng Cheng 0005, Rathijit Sen, Matteo Interlandi |
DaMoN | 8 |
| 2023 | ARK: GPU-driven Code Execution for Distributed Deep Learning
Changho Hwang, KyoungSoo Park, Ran Shu 0001, Xinyuan Qu, Peng Cheng 0005, Yongqiang Xiong |
NSDI | 5 |
| 2023 | Poster: Meili: Towards SmartNIC as a ServiceabstractThe gap between the stagnation of CPU power and the increase in network bandwidth has promoted a shift towards placing more computation on network hardware [16, 17]. Therefore, SmartNICs have become prevalent in data centers to serve various cloud applications, from network functions [15, 17, 22] to high-level applications like distributed applications and storage [14, 16, 18--21, 23]. Shaofeng Wu, Zhixiong Niu, Ran Shu 0001, Peng Cheng 0005, Yongqiang Xiong, Chun Jason Xue, Zaoxing Liu, Hong Xu 0001 |
SIGCOMM | 5 |
| 2023 | SPFresh: Incremental In-Place Update for Billion-Scale Vector SearchabstractApproximate Nearest Neighbor Search (ANNS) on high dimensional vector data is now widely used in various applications, including information retrieval, question answering, and recommendation. As the amount of vector data grows continuously, it becomes important to support updates to vector index, the enabling technique that allows for efficient and accurate ANNS on vectors. Yuming Xu, Hengyu Liang, Jin Li 0050, Shuotao Xu, Qi Chen 0009, Qianxi Zhang, Cheng Li 0001, Ziyue Yang 0002, Fan Yang 0024, Yuqing Yang 0001, Peng Cheng 0005, Mao Yang 0004 |
SOSP | 11 |
| 2022 | OpenNetLab: Open Platform for RL-based Congestion Control for Real-Time CommunicationsabstractWith the growing importance of real-time communications (RTC), designing congestion control (CC) algorithms for RTC that achieve high network performance and QoE is gaining attention. Recently, data-driven, reinforcement learning (RL)-based CC algorithms for RTC have shown great potential, outperforming traditional rule-based counterparts. However, there are no open platforms tailored for training, evaluation, and validation of the algorithms that can facilitate this emerging research area. Jeongyoon Eo, Zhixiong Niu, Wenxue Cheng, Francis Y. Yan, Jorina Kardhashi, Scott Inglis, Michael Revow, Byung-Gon Chun, Peng Cheng 0005, Yongqiang Xiong |
APNet | 10 |
| 2022 | A Disaggregate Data Collecting Approach for Loss-Tolerant ApplicationsabstractDatacenter generates operation data at an extremely high rate, and data center operators collect and analyze them for problem diagnosis, resource utilization improvement, and performance optimization. However, existing data collection methods fail to efficiently aggregate and store data at extremely high speed and scale. In this paper, we explore a new approach that leverages programmable switches to aggregate data and directly write data to the destination storage. Our proposed data collection system, ALT, uses programmable switches to control NVMe SSDs on remote hosts without the involvement of a remote CPU. To tolerate loss, ALT uses an elegant data structure to enable efficient data recovery when retrieving the collected data. We implement our system on a Tofino-based programmable switch for a prototype. Our evaluation shows that ALT can saturate SSD’s peak performance without any CPU involvement. Ziyuan Liu 0008, Zhixiong Niu, Ran Shu 0001, Wenxue Cheng, Peng Cheng 0005, Yongqiang Xiong, Jacob Nelson 0001, Dan R. K. Ports |
APNet | 5 |
| 2022 | PipeDevice: a hardware-software co-design approach to intra-host container communicationabstractContainers are prevalently adopted due to the deployment and performance advantages over virtual machines. For many containerized data-intensive applications, however, bulky data transfers may pose performance issues. In particular, communication across co-located containers on the same host incurs large overheads in memory copy and the kernel's TCP stack. Existing solutions such as shared-memory networking and RDMA have their own limitations, including insufficient memory isolation and limited scalability. Chuanwen Wang, Zhixiong Niu, Ran Shu 0001, Peng Cheng 0005, Yongqiang Xiong, Dongsu Han, Chun Jason Xue, Hong Xu 0001 |
CoNEXT | 5 |
| 2022 | Moneo: Non-intrusive Fine-grained Monitor for AI InfrastructureabstractCloud-based AI infrastructure is increasingly important, especially on large-scale distributed training. To improve its efficiency and serviceability, real-time monitoring of the infrastructure and profiling the workload are proved to be the effective approach empirically. However, cloud environment poses great challenges as service providers cannot interfere with their tenants' workloads or touch user data, thus previous instrumentation-based monitoring approach cannot be applied, nor does the workload trace collection.We propose Moneo, a non-intrusive cloud-friendly monitoring system for AI infrastructure. Moneo is capable of intelligently collecting the key architecture-level metrics at finer granularity in real-time without instrumenting or tracing the workloads, which has been deployed in real production cloud, Azure. We analyze the results reported by Moneo for typical large-scale distributed AI workloads from real deployment. Results demonstrate that Moneo can effectively help service providers understand the real resource usage patterns of various AI workloads and real networking requirements, so as to get valuable findings help improve the efficiency of cloud infrastructure and optimize the software stack with the consideration of the characteristic resource usage requirements for different AI workloads. Yifan Xiong 0001, Chen Tian 0001, Peng Cheng 0005, Yongqiang Xiong |
ICC | 6 |
| 2022 | RuleCache: Accelerating Web Application Firewalls by On-line Learning Traffic PatternsabstractWeb Application Firewall (WAF) is widely deployed in cloud to protect web applications, whose performance becomes one of the major bottlenecks for web services. In this paper, we comprehensively analyze several root causes that downgrade WAF’s efficiency. Inspired by that, we build a caching system RuleCache to devise optimization strategies for improving WAF’s performance. Among, Rule Ordering Cache is online learning an optimal order of the ruleset for a better performance of blocking. Rule Result Cache reuses rule results of targets, saving large repetitive computations. Additionally, Rule Prepruning Cache aims to cut extra overhead by processing the static rules in the offline stage. Our evaluation demonstrates that the prototype can improve the performance by up to 3.85x, 1.57x, and 2.4x respectively with the above modules, and up to 5.5x in total. Qingni Shen, Peng Cheng 0005, Yongqiang Xiong, Zhonghai Wu |
ICWS | 3 |
| 2022 | An Adaptive Deep RL Method for Non-Stationary Environments with Piecewise Stable ContextabstractOne of the key challenges in deploying RL to real-world applications is to adapt to variations of unknown environment contexts, such as changing terrains in robotic tasks and fluctuated bandwidth in congestion control. Existing works on adaptation to unknown environment contexts either assume the contexts are the same for the whole episode or assume the context variables are Markovian. However, in many real-world applications, the environment context usually stays stable for a stochastic period and then changes in an abrupt and unpredictable manner within an episode, resulting in a segment structure, which existing works fail to address. To leverage the segment structure of piecewise stable context in real-world applications, in this paper, we propose a \textit{\textbf{Se}gmented \textbf{C}ontext \textbf{B}elief \textbf{A}ugmented \textbf{D}eep~(SeCBAD)} RL method. Our method can jointly infer the belief distribution over latent context with the posterior over segment length and perform more accurate belief context inference with observed data within the current context segment. The inferred belief context can be leveraged to augment the state, leading to a policy that can adapt to abrupt variations in context. We demonstrate empirically that SeCBAD can infer context segment length accurately and outperform existing methods on a toy grid world environment and Mujuco tasks with piecewise-stable context. Xiangming Zhu 0002, Pushi Zhang, Li Zhao 0007, Wenxue Cheng, Peng Cheng 0005, Yongqiang Xiong, Tao Qin 0001, Jianyu Chen 0002, Tie-Yan Liu |
NeurIPS | 7 |
| 2022 | PilotFish: Harvesting Free Cycles of Cloud Gaming with Deep Learning Training
Wei Zhang 0149, Binghao Chen, Zhenhua Han, Quan Chen 0002, Peng Cheng 0005, Fan Yang 0024, Ran Shu 0001, Yuqing Yang 0001, Minyi Guo |
USENIX ATC | 5 |
| 2022 | NetKernel: Making Network Stack Part of the Virtualized InfrastructureabstractThis paper presents a system called NetKernel that decouples the network stack from the guest virtual machine and offers it as an independent module. NetKernel represents a new paradigm where network stack can be managed as part of the virtualized infrastructure. It provides important efficiency benefits: By gaining control and visibility of the network stack, operators can perform network management more directly and flexibly, such as multiplexing VMs running different applications to the same network stack module to save CPU cores, and enforcing fair bandwidth sharing. Users also benefit from the simplified stack deployment and better performance: For example mTCP can be deployed without API change to support nginx natively, and shared memory networking can be readily enabled to improve performance of colocated VMs. Testbed evaluation using 100G NICs shows that NetKernel preserves the performance and scalability of both kernel and userspace network stacks, and provides the same isolation as the current architecture. Zhixiong Niu, Peng Cheng 0005, Yongqiang Xiong, Dongsu Han, Keith Winstein, Chun Jason Xue, Hong Xu 0001 |
IEEE/ACM Trans. Netw. | 3 |
| 2021 | NFD: Using Behavior Models to Develop Cross-Platform Network FunctionsabstractNFV ecosystem is flourishing and more and more NF platforms appear, but this makes NF vendors difficult to deliver NFs rapidly to diverse platforms. We propose an NF development framework named NFD for cross-platform NF development. NFD's main idea is to decouple the functional logic from the platform logic -it provides a platform-independent language to program NFs' behavior models, and a compiler with interfaces to develop platform-specific plugins. By enabling a plugin on the compiler, various NF models would be compiled to executables integrated with the target platform. We prototype NFD, build 14 NFs, and support 6 platforms (standard Linux, OpenNetVM, GPU, SGX, DPDK, OpenNF). Our evaluation shows that NFD can save development workload for cross-platform NFs and output valid and performant NFs. Hongyi Huang, Wenfei Wu, Yongchao He, Bangwen Deng, Ying Zhang 0022, Yongqiang Xiong, Guo Chen 0001, Yong Cui 0001, Peng Cheng 0005 |
INFOCOM | 9 |
| 2020 | NetKernel: Making Network Stack Part of the Virtualized Infrastructure
Zhixiong Niu, Hong Xu 0001, Peng Cheng 0005, Yongqiang Xiong, Tao Wang 0088, Dongsu Han, Keith Winstein |
USENIX ATC | 3 |
| 2020 | Observing and Mitigating Micro-Burst Traffic in Data Center NetworksabstractMicro-burst traffic is not uncommon in data centers. It can cause packet dropping, which may result in serious performance degradation (e.g., Incast problem). However, current approaches to mitigate micro-burst is usually ad-hoc and not based on a principled understanding of the underlying behaviors. On the other hand, traditional studies focus on traffic burstiness in a single flow, while micro-burst traffic in the data centers could occur with highly fan-in communication pattern, and its dynamic behavior is still unclear. To this end, in this paper, we re-examine the micro-burst traffic in typical data center scenarios. We find that the evolution of micro-burst is determined by both TCP's self-clocking mechanism and congestion control algorithm. Besides, dynamic behaviors of micro-burst under various scenarios can all be described by the time derivative of queue length evolution.Our observations also implicate that conventional solutions like absorbing and pacing are ineffective to mitigate micro-burst traffic.Instead, senders need to rapidly respond to some explicit signals of the queue buildup caused by the micro-burst traffic rather than independently and ineffectually pacing themselves in isolation. Inspired by the findings and insights from experimental observations, we propose Micro-burst-Aware Transport Control Protocol (MATCP), which leverages characteristic behaviors of micro-burst traffic derived from the time derivative of the queue occupancy. MATCP can suppress the sharp queue length increment by over 2x and reduce the tail query completion time by up to 84.4%. Danfeng Shan, Fengyuan Ren, Peng Cheng 0005, Ran Shu 0001, Chuanxiong Guo |
IEEE/ACM Trans. Netw. | 3 |
| 2019 | DLBooster: Boosting End-to-End Deep Learning Workflows with Offloading Data Preprocessing PipelinesabstractIn recent years, deep learning (DL) has prospered again due to improvements in both computing and learning theory. Emerging studies mostly focus on the acceleration of refining DL models but ignore data preprocessing issues. However, data preprocessing can significantly affect the overall performance of end-to-end DL workflows. Our studies on several image DL workloads show that existing preprocessing backends are quite inefficient: they either perform poorly in throughput (30% degradation) or burn too many (>10) CPU cores. Based on these observations, we propose DLBooster, a high-performance data preprocessing pipeline that selectively offloads key workloads to FPGAs, to fit the stringent demands on data preprocessing for cutting-edge DL applications. Our testbed experiments show that, compared with the existing baselines, DLBooster can achieve 1.35×~2.4× image processing throughput in several DL workloads, but consumes only 1/10 CPU cores. Besides, it also reduces the latency by 1/3 in online image inference. Dan Li 0001, Binyao Jiang, Xi Fan, Jinkun Geng, Wei Bai 0001, Ran Shu 0001, Peng Cheng 0005, Yongqiang Xiong |
ICPP | 12 |
| 2019 | Direct Universal Access: Making Data Center Resources Available to FPGA
Ran Shu 0001, Peng Cheng 0005, Guo Chen 0001, Yongqiang Xiong, Derek Chiou, Thomas Moscibroda |
NSDI | 2 |
| 2019 | MP-RDMA: Enabling RDMA With Multi-Path Transport in DatacentersabstractRDMA is becoming prevalent because of its low latency, high throughput and low CPU overhead. However, in current datacenters, RDMA remains a single path transport which is prone to failures and falls short to utilize the rich parallel network paths. Unlike previous multi-path approaches, which mainly focus on TCP, this paper presents a multi-path transport for RDMA, i.e. MP-RDMA, which efficiently utilizes the rich network paths in datacenters. MP-RDMA employs three novel techniques to address the challenge of limited RDMA NICs on-chip memory size: 1) a multi-path ACK-clocking mechanism to distribute traffic in a congestion-aware manner without incurring per-path states; 2) an out-of-order aware path selection mechanism to control the level of out-of-order delivered packets, thus minimizes the meta data required to them; 3) a synchronise mechanism to ensure in-order memory update whenever needed. With all these techniques, MP-RDMA only adds 66B to each connection state compared to single-path RDMA. Our evaluation with an FPGA-based prototype demonstrates that compared with single-path RDMA, MP-RDMA can significantly improve the robustness under failures ( $2\times \sim 4\times $ higher throughput under 0.5%~10% link loss ratio) and improve the overall network utilization by up to 47%. Guo Chen 0001, Yuanwei Lu, Bojie Li, Kun Tan 0002, Yongqiang Xiong, Peng Cheng 0005, Jiansong Zhang 0001, Thomas Moscibroda |
IEEE/ACM Trans. Netw. | 6 |
| 2019 | Tagger: Practical PFC Deadlock Prevention in Data Center NetworksabstractRemote direct memory access over converged Ethernet deployments is vulnerable to deadlocks induced by priority flow control. Prior solutions for deadlock prevention either require significant changes to routing protocols or require excessive buffers in the switches. In this paper, we propose Tagger, a scheme for deadlock prevention. It does not require any changes to the routing protocol and needs only modest buffers. Tagger is based on the insight that given a set of expected lossless routes, a simple tagging scheme can be developed to ensure that no deadlock will occur under any failure conditions. Packets that do not travel on these lossless routes may be dropped under extreme conditions. We design such a scheme, prove that it prevents deadlock, and implement it efficiently on commodity hardware. Shuihai Hu, Yibo Zhu 0001, Peng Cheng 0005, Chuanxiong Guo, Jitendra Padhye, Kai Chen 0005 |
IEEE/ACM Trans. Netw. | 3 |
| 2018 | Micro-Burst in Data Centers: Observations, Analysis, and MitigationsabstractMicro-burst traffic is not uncommon in data centers. It can cause packet dropping, which results in serious performance degradation (e.g., Incast problem). However, current solutions that attempt to suppress micro-burst traffic are extrinsic and ad hoc, since they lack the comprehensive and essential understanding of micro-burst's root cause and dynamic behavior. On the other hand, traditional studies focus on traffic burstiness in a single flow, while in data centers micro-burst traffic could occur with highly fan-in communication pattern, and its dynamic behavior is still unclear. To this end, in this paper, we re-examine the microburst traffic in typical data center scenarios. We find that evolution of micro-burst is determined by both TCP's self-clocking mechanism and bottleneck link. Besides, dynamic behaviors of micro-burst under various scenarios can all be described by the slope of queue length evolution. Our observations also implicate that conventional solutions like absorbing and pacing are ineffective to mitigate micro-burst traffic. Instead, senders need to slow down as soon as possible. Inspired by the findings and insights from experimental observations, we propose S-ECN policy, which is an ECN marking policy leveraging the slope of queue length evolution. Transport protocols utilizing S-ECN policy can suppress the sharp queue length increment by over 2×, and reduce the average query completion time by ~12-27%. Danfeng Shan, Fengyuan Ren, Peng Cheng 0005, Ran Shu 0001, Chuanxiong Guo |
ICNP | 3 |
| 2018 | Multi-Path Transport for RDMA in Datacenters
Yuanwei Lu, Guo Chen 0001, Bojie Li, Kun Tan 0002, Yongqiang Xiong, Peng Cheng 0005, Jiansong Zhang 0001, Enhong Chen, Thomas Moscibroda |
NSDI | 6 |
| 2018 | FUSO: Fast Multi-Path Loss Recovery for Data Center Networks
Guo Chen 0001, Yuanwei Lu, Yuan Meng 0002, Bojie Li, Kun Tan 0002, Dan Pei, Peng Cheng 0005, Layong Luo, Yongqiang Xiong, Xiaoliang Wang 0001, Youjian Zhao |
IEEE/ACM Trans. Netw. | 7 |
| 2017 | Memory Efficient Loss Recovery for Hardware-based Transport in DatacenterabstractLimited by the small on-chip memory, hardware-based transport typically implements go-back-N loss recovery mechanism, which costs very few memory but is well-known to perform inferior even under small packet loss ratio. We present MELO, an efficient selective retransmission mechanism for hardware-based transport, which consumes only a constant small memory regardless of the number of concurrent connections. Specifically, MELO employs an architectural separation between data and meta data storage and uses a shared bits pool allocation mechanism to reduce meta data on-chip memory footprint. By only adding in average 23B extra on-chip states for each connection, MELO achieves up to 14.02x throughput while reduces 99% tail FCT by 3.11x compared with go-back-N under certain loss ratio. Yuanwei Lu, Guo Chen 0001, Zhenyuan Ruan, Wencong Xiao, Bojie Li, Jiansong Zhang 0001, Yongqiang Xiong, Peng Cheng 0005, Enhong Chen |
APNet | 8 |
| 2017 | Tagger: Practical PFC Deadlock Prevention in Data Center NetworksabstractRemote Direct Memory Access over Converged Ethernet (RoCE) deployments are vulnerable to deadlocks induced by Priority Flow Control (PFC). Prior solutions for deadlock prevention either require signi.cant changes to routing protocols, or require excessive bu.ers in the switches. In this paper, we propose Tagger, a scheme for deadlock prevention. It does not require any changes to the routing protocol, and needs only modest bu.ers. Tagger is based on the insight that given a set of expected lossless routes, a simple tagging scheme can be developed to ensure that no deadlock will occur under any failure conditions. Packets that do not travel on these lossless routes may be dropped under extreme conditions. We design such a scheme, prove that it prevents deadlock and implement it e.ciently on commodity hardware. Shuihai Hu, Peng Cheng 0005, Chuanxiong Guo, Jitendra Padhye, Kai Chen 0005 |
CoNEXT | 3 |
| 2017 | Network Stack as a Service in the CloudabstractThe tenant network stack is implemented inside the virtual machines in today's public cloud. This legacy architecture presents a barrier to protocol stack innovation due to the tight coupling between the network stack and the guest OS. In particular, it causes many deployment troubles to tenants and management and efficiency problems to the cloud provider. To address these issues, we articulate a vision of providing the network stack as a service. The central idea is to decouple the network stack from the guest OS, and offer it as an independent entity implemented by the cloud provider. This re-architecting allows tenants to readily deploy any stack independent of its kernel, and the provider to offer meaningful SLAs to tenants by gaining control over the network stack. We sketch an initial design called NetKernel to accomplish this vision. Our preliminary testbed evaluation with a prototype shows the feasibility and benefits of our idea. Zhixiong Niu, Hong Xu 0001, Dongsu Han, Peng Cheng 0005, Yongqiang Xiong, Guo Chen 0001, Keith Winstein |
HotNets | 4 |
| 2017 | Performance analysis of randomized data fetching in cluster computingabstractThe shuffle transfer pattern is widely adopted in today's cluster computing applications and the completion time of each group of transmissions directly affects application performance. Because of the restriction on the number of concurrent threads and the TCP Incast problem, the randomized data fetching strategy is widely employed in this kind of communication in practice. In this paper, to assess the performance of randomized data fetching, we build a general analytical model and define two metrics - link overload probability and K-deviation load balancing probability - to evaluate the degree of link overload and load balancing respectively, since they are closely related to the transfer completion time. Leveraging our model, we theoretically analyze the transfer performance in three typical scenarios and provide recommendations for setting the number of concurrent connections per receiver. Finally, we validate the theoretical analysis as well as the recommendations through extensive simulations. Tong Zhang 0018, Peng Cheng 0005, Wenxue Cheng, Bo Wang 0066, Fengyuan Ren |
IWQoS | 2 |
| 2016 | TFC: token flow control in data center networksabstractServices in modern data center networks pose growing performance demands. However, the widely existed special traffic patterns, such as micro-burst, highly concurrent flows, on-off pattern of flow transmission, exacerbate the performance of transport protocols. In this work, an clean-slate explicit transport control mechanism, called Token Flow Control (TFC), is proposed for data center networks to achieve high link utilization, ultra-low latency, fast convergence, and rare packets dropping. TFC uses tokens to represent the link bandwidth resource and define the concept of effective flows to stand for consumers. The total tokens will be explicitly allocated to each consumer every time slot. TFC excludes in-network buffer space from the flow pipeline and thus achieves zero-queueing. Besides, a packet delay function is added at switches to prevent packets dropping with highly concurrent flows. The performance of TFC is evaluated using both experiments on a small real testbed and large-scale simulations. The results show that TFC achieves high throughput, fast convergence, near zero-queuing and rare packets loss in various scenarios. Jiao Zhang 0002, Fengyuan Ren, Ran Shu 0001, Peng Cheng 0005 |
EuroSys | 4 |
| 2016 | Deadlocks in Datacenter Networks: Why Do They Form, and How to Avoid ThemabstractDriven by the need for ultra-low latency, high throughput and low CPU overhead, Remote Direct Memory Access (RDMA) is being deployed by many cloud providers. To deploy RDMA in Ethernet networks, Priority-based Flow Control (PFC) must be used. PFC, however, makes Ethernet networks prone to deadlocks. Prior work on deadlock avoidance has focused on {\em necessary} condition for deadlock formation, which leads to rather onerous and expensive solutions for deadlock avoidance. In this paper, we investigate {\em sufficient} conditions for deadlock formation, conjecturing that avoiding {\em sufficient} conditions might be less onerous. Shuihai Hu, Peng Cheng 0005, Chuanxiong Guo, Jitendra Padhye, Kai Chen 0005 |
HotNets | 3 |
| 2016 | ClickNP: Highly flexible and High-performance Network Processing with Reconfigurable HardwareabstractHighly flexible software network functions (NFs) are crucial components to enable multi-tenancy in the clouds. However, software packet processing on a commodity server has limited capacity and induces high latency. While software NFs could scale out using more servers, doing so adds significant cost. This paper focuses on accelerating NFs with programmable hardware, i.e., FPGA, which is now a mature technology and inexpensive for datacenters. However, FPGA is predominately programmed using low-level hardware description languages (HDLs), which are hard to code and difficult to debug. More importantly, HDLs are almost inaccessible for most software programmers. This paper presents ClickNP, a FPGA-accelerated platform for highly flexible and high-performance NFs with commodity servers. ClickNP is highly flexible as it is completely programmable using high-level C-like languages, and exposes a modular programming abstraction that resembles Click Modular Router. ClickNP is also high performance. Our prototype NFs show that they can process traffic at up to 200 million packets per second with ultra-low latency ($< 2\mu$s). Compared to existing software counterparts, with FPGA, ClickNP improves throughput by 10x, while reducing latency by 10x. To the best of our knowledge, ClickNP is the first FPGA-accelerated platform for NFs, written completely in high-level language and achieving 40 Gbps line rate at any packet size. Bojie Li, Kun Tan 0002, Layong Luo, Yanqing Peng, Renqian Luo, Ningyi Xu, Yongqiang Xiong, Peng Cheng 0005 |
SIGCOMM | 8 |
| 2016 | Fast and Cautious: Leveraging Multi-path Diversity for Transport Loss Recovery in Data Centers
Guo Chen 0001, Yuanwei Lu, Yuan Meng 0002, Bojie Li, Kun Tan 0002, Dan Pei, Peng Cheng 0005, Layong Luo, Yongqiang Xiong, Xiaoliang Wang 0001, Youjian Zhao |
USENIX ATC | 7 |
| 2016 | An Energy Efficiency Perspective on Rate Adaptation for 802.11n NICabstractRate adaptation (RA) has been traditionally used to achieve high goodput. In this work, we design RA for 802.11n NICs from an energy-efficiency perspective. We show that current MIMO RA algorithms are not energy efficient for NICs despite ensuring high throughput. The fundamental problem is that, the high-throughput setting is not equivalent to the energy-efficient one. Marginal throughput gain may be realized at high energy cost. We then propose EERA and EERA+, two energy-based RA schemes that trade off goodput for energy savings at NICs. EERA applies multidimensional ternary search and simultaneous pruning to speed up its runtime convergence in single-client operations, and uses fair airtime sharing to handle multiple-client operations. EERA+ further searches for multiple, staged rates to yield more energy savings over EERA. Our experiments have confirmed their effectiveness in various scenarios. Chi-Yu Li 0001, Chunyi Peng 0001, Peng Cheng 0005, Songwu Lu, Xinbing Wang, Fengyuan Ren, Tao Wang 0004 |
IEEE Trans. Mob. Comput. | 3 |
| 2015 | Slowing Little Quickens More: Improving DCTCP for Massive Concurrent FlowsabstractDCTCP is a potential TCP replacement to satisfy the requirements of data center network. It receives wide concerns in both academic and industrial circles. However, DCTCP could only support tens of concurrent flows well and suffers timeouts and throughput collapse facing numerous concurrent flows. This is far from the requirement of data center network. Data centers employing partition/aggregation pattern usually involve hundreds of concurrent flows. In this paper, after tracing DCTCP's dynamic behavior through experiments, we explored two roots for DCTCP's failure under the high fan-in traffic pattern: (1) The regulation mechanism of sending window is ineffective when cwnd is decreased to the minimum size, (2) The bursts induced by synchronized flows with small cwnd cause fatal packet loss leading to severe timeouts. We enhance DCTCP to support massive concurrent flows by regulating the sending time interval and desynchronizing the sending time in particular conditions. The new protocol called DCTCP+ outperforms DCTCP when the number of concurrent flows increases to several hundreds. DCTCP+ can normally work to effectively support the short concurrent query responses in the benchmark from real production clusters, and keep the same good performance with the mixture of background traffic. Mao Miao, Peng Cheng 0005, Fengyuan Ren, Ran Shu 0001 |
ICPP | 2 |
| 2014 | Catch the Whole Lot in an Action: Rapid Precise Packet Loss Notification in Data Center
Peng Cheng 0005, Fengyuan Ren, Ran Shu 0001, Chuang Lin 0002 |
NSDI | 1 |
| 2013 | Ease the Queue Oscillation: Analysis and Enhancement of DCTCPabstractBecause of the terrible performance of TCP protocol in data center environment, DCTCP has been proposed as a TCP replacement, which uses a simple marking mechanism at switches and a few amendments at end hosts to adjust congestion window based on the extent of the congestion in networks. Thus, DCTCP can make a proper tradeoff between high throughput and low latency. However, through our observation, we discover that DCTCP causes severe oscillation of queue under some parameters and network configuration. Our perceptual analysis concludes that the rough single-threshold marking mechanism may be the essential reason. Therefore, we propose Double-Threshold DCTCP as an improvement of DCTCP. Then, by applying describing function method in nonlinear control theory, we analyze the stability of both DCTCP and Double-Threshold DCTCP, and theoretically explain why Double-Threshold DCTCP is more stable than DCTCP. At last, we validate theoretical analysis and conclude that the Double- Threshold DCTCP can achieve smaller queue, and the queue length of Double-Threshold DCTCP is less sensitive to the growing number of flows. Further, Double-Threshold DCTCP can postpone the throughput collapse caused by Incast traffic and reduce the tail latency in completion time experiment. Wen Chen 0026, Peng Cheng 0005, Fengyuan Ren, Ran Shu 0001, Chuang Lin 0002 |
ICDCS | 2 |