VLDB 2026 Research / reviewers in the wild / expert
Yongqiang Xiong
dblp:57/4569
· DBLP profile ↗
94ranked-venue papers
0as first author
42since 2021 · last 2026
0000-0003-4175-0097ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 59 · 30 since 2021Systems, architecture and hardware · 21 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5Software engineering, systems software and programming languages · 4 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Learned Switch Behavior Modeling
Daqian Ding, Zhixiong Niu, Ziyu Mao, Jingyu Wang 0001, Yongqiang Xiong, Yiming Qiu 0001 |
APNet | 6 |
| 2026 | Saving Lost Intents in Network ConfigurationsabstractNetwork configurations in brownfield environments often degenerate into “append-only” artifacts whose original design intent has been lost to personnel churn and documentation decay. In this paper, we identify intent recovery as a distinct problem: given an observed network state of interest, the goal is to infer why it exists rather than merely what it does. We characterize this task as an ill-posed inverse problem, because the compilation from high-level intent to low-level configuration discards rich contextual information, making the inverse mapping inherently one-to-many. To resolve this ambiguity, we observe that the original design process leaves residual traces in surrounding operational artifacts. We propose Matt, a framework that regularizes the inference by mining and fusing three complementary dimensions of such residual context: Semantics (business hierarchy), Provenance (configuration structure), and History (evolutionary timeline). An ablation study on synthetic datasets calibrated to real campus features confirms that each dimension provides a unique, irreplaceable contribution, with the full pipeline achieving 0.71 Exact Match accuracy and 0.79 Intent F1 score. Zhixiong Niu, Yongqiang Xiong, Hong Xu 0001 |
APNet | 4 |
| 2026 | OptiFlow: Towards LLM-Driven Optimization of Collective Communication Algorithms
Ziyue Yang 0002, Kaihui Gao, Shuai Wang 0028, Li Chen 0008, Zhixiong Niu, Ran Shu 0001, Wenxue Cheng, Peng Cheng 0005, Yongqiang Xiong, Dan Li 0001 |
APNet | 10 |
| 2026 | PIN: Less Is More for RDMA Load BalancingabstractAs RDMA becomes increasingly tolerant to out-of-order delivery, fine-grained packet-level multipathing is emerging as a practical design for datacenter fabrics. Packet spraying and related schemes improve load distribution for large flows, but they also force RDMA flows onto multiple paths whose conditions can differ at short timescales due to randomized traffic placement. For short flows, even one packet sent on a temporarily slower path can delay the entire flow. As a result, fine-grained load balancing can hurt, rather than help, small multi-packet flows. Jichun Wu, Ran Shu 0001, Yongqiang Xiong |
APNet | 4 |
| 2026 | Offloading Cloud Network Services at Production Scale with SONiC DASH SmartSwitch
Shaofeng Wu, Zhixiong Niu, Riff Jiang, Lawrence Lee, Junhua Zhai, Ze Gan, Vasundhara Volam, Prabhat Aravind, Prince Sunny, Prince George, Evan Langlais, Soumya Tiwari, Venkat Satish Katta, Weixi Chen, Rishiraj Hazarika, Sachin Jain, Deven Jagasia, Michal Zygmunt, Avijit Gupta, Neeraj Motwani, Pranjal Shrivastava, Anil Reddy Pannala, Kristina Moore, James Grantham, Anupam Pandey, Guohan Lu, Gerald DeGrace, Rishabh Tewari, Erica Lan, Deepak Bansal, David A. Maltz, Yongqiang Xiong, Hong Xu 0001 |
NSDI | 36 |
| 2026 | SmartNIC-Enabled Live Migration for Storage-Optimized VMs with PYROCUMULUS
Jiechen Zhao 0002, Ran Shu 0001, Ziyue Yang 0002, Rui Ma 0021, Derek Chiou, Natalie D. Enright Jerger, Peng Cheng 0005, Yongqiang Xiong |
NSDI | 9 |
| 2026 | Nüwa: A Generative Control Plane for AI Network SimulationabstractNetwork simulation plays a critical role in improving the efficiency of large-scale AI clusters for design validation, parameter tuning, and protocol development. However, high-fidelity network simulation becomes prohibitively slow at scale, especially when running large batches of experiments on topologies with tens or hundreds of thousands of accelerators. We observe that a key bottleneck comes from the control plane. Existing network simulators typically compute routes and install forwarding tables at initialization, which can consume hundreds of GB of memory before packet-event execution begins and limit overall simulation throughput. In this paper, we present Nüwa, which views routing as a compilation problem, it leverages the hierarchical and symmetric structure common in AI fabrics and compiles a declarative topology description together with routing policies into compact forwarding artifacts that are fast to generate and efficient to look up. Evaluations show that Nüwa can reduce simulation initialization time from hours to only 25 seconds for a 65,536-GPU cluster. For end-to-end simulation time, Nüwa takes only 20% of that required by existing approaches in a 40K+ GPU cluster, and Nüwa can scale to a 221,184-GPU cluster. Ran Shu 0001, Peng Zhang 0011, Danfeng Shan, Yongqiang Xiong |
SIGCOMM | 6 |
| 2026 | STORM: Enabling Traffic Scheduling for RDMAabstractRemote Direct Memory Access (RDMA) is increasingly used as a shared communication substrate across datacenter workloads with very different scheduling needs, from request-response services and storage fan-out to AI training collectives. Proper request scheduling can reduce communication time, but in practice, no RDMA flow scheduling is enabled in datacenters, leaving traffic to simple fair sharing. We present STORM, a NIC-level scheduler for all types of RDMA workloads using NIC-only information: the known RDMA request size, and per-queue-pair backlog. STORM converts these signals into a small number of extra priority levels on the wire and prioritizes requests that are either near completion or blocking queued dependent work. STORM requires no application hints and works with both in-order RoCEv2 and newer RDMA stacks that tolerate reordering. We prototype STORM on an FPGA NIC with negligible overhead. Across representative cloud and LLM training workloads, STORM reduces training iteration time by up to 12% and reduces average and P99 flow completion slowdown by up to 90% compared to fair scheduling. Jichun Wu, Ran Shu 0001, Gianni Antichi, Yongqiang Xiong, Jon Crowcroft |
SIGCOMM | 4 |
| 2026 | SuperBench: A Proactive Validation System for Improving Reliability of Cloud AI InfrastructureabstractReliability in cloud AI infrastructure is crucial for cloud service providers, prompting the widespread use of hardware redundancies. However, these redundancies can inadvertently lead to hidden degradation, known as “gray failure”, for AI workloads, significantly affecting end-to-end performance and concealing performance issues, which complicates root cause analysis for failures and regressions. We introduce SuperBench, a proactive validation system for AI infrastructure that mitigates hidden degradation caused by hardware redundancies and enhances overall reliability. SuperBench features a comprehensive benchmark suite, capable of evaluating individual hardware components and representing most real AI workloads. It comprises a Validator that learns benchmark criteria to pinpoint defective components clearly. Additionally, SuperBench incorporates a Selector to balance validation time and issue-related penalties, enabling optimal timing for validation execution with a tailored subset of benchmarks. Through testbed evaluation and simulation, we demonstrate that SuperBench can increase the mean time between incidents by up to 22.61×. SuperBench has been successfully deployed in Azure production, validating hundreds of thousands of GPUs every year. Yifan Xiong 0001, Ziyue Yang 0002, Guoshuai Zhao 0001, Dong Zhong, Boris Pinzur, Jie Zhang 0048, Yang Wang 0053, Hossein Pourreza, Jeff Baxter, Kushal Datta, Prabhat Ram, Luke Melton, Joe Chau, Peng Cheng 0005, Yongqiang Xiong, Lidong Zhou |
ACM Trans. Comput. Syst. | 19 |
| 2026 | Fine-Grained Scheduling of In-Network Aggregation Resources for Efficient Machine Learning Service
Shichen Dong, Zhixiong Niu, Mingchao Zhang, Zhiying Xu, Chuntao Hu, Pengzhi Zhu, Qingchun Song, Peng Cheng 0005, Cam-Tu Nguyen, Shaoling Sun, Xiaohu Xu, Yongqiang Xiong, Wei Wang 0002, Xiaoliang Wang 0001, Guihai Chen |
IEEE Trans. Netw. | 13 |
| 2025 | Nüwa: Efficient Generative Control Plane for AI Network Simulation
Ran Shu 0001, Peng Zhang 0011, Yongqiang Xiong |
APNet | 4 |
| 2025 | Miniature: Fast AI Supercomputer Networks Simulation on FPGAs
Yicheng Qian, Ran Shu 0001, Rui Ma 0021, Yang Wang 0053, Derek Chiou, Nadeen Gebara, Luca Piccolboni, Miriam Leeser, Yongqiang Xiong |
APNet | 9 |
| 2025 | SRC: A Scalable Reliable Connection for RDMA with Decoupled QPs and Connections
Ran Shu 0001, Yongqiang Xiong |
APNet | 3 |
| 2025 | Daredevil: Rescue Your Flash Storage from Inflexible Kernel Storage StackabstractExisting kernel storage stacks for NVMe SSDs struggle to address performance interference between I/O requests from tenants with different SLAs, leading to the multi-tenancy issue. Addressing this requires separating their I/O requests within the NVMe I/O queues (NQs). However, our analysis reveals that the static CPU core-NQ bindings of current storage stacks restrict their flexibility to achieve this goal. Ran Shu 0001, Jiayi Lin 0007, Qingyu Zhang 0005, Ziyue Yang 0002, Jie Zhang 0048, Yongqiang Xiong, Chenxiong Qian |
EuroSys | 7 |
| 2025 | 3L-Cache: Low Overhead and Precise Learning-based Eviction Policy for Caches
Zhixiong Niu, Yongqiang Xiong |
FAST | 3 |
| 2025 | NetSophon: Enabling Runtime Copilot for Programmable Dataplane for Cloud OperatorsabstractRuntime traffic analysis on programmable data-plane requires substantial human effort, and the high speed and complexity of dataplane often make human capacity the efficiency bottleneck. While existing work has proposed LLM-based approaches, they typically rely on offline network logs, failing to address the human capacity limitations in real-time environments. This paper explores the potential of leveraging evolving LLMs to mitigate these human-centric challenges in real physical dataplane. It outlines a novel framework called NetSophon, which features an LLM-based brain for efficient decision-making and an effective arm to manipulate and perceive the physical programmable dataplane. Through interactions among the brain, arm, and dataplane, NetSophon acts as a "super-copilot" for human operators, facilitating real-time dataplane traffic analysis at scale. A case study demonstrates NetSophon’s potential to assist human operators in interacting with dataplane. Shaofeng Wu, Zhixiong Niu, Riff Jiang, Lizhao You, Qiao Xiang, Hong Xu 0001, Yongqiang Xiong |
ICNP | 9 |
| 2025 | Mina: Fine-Grained In-network Aggregation Resource Scheduling for Machine Learning Service
Shichen Dong, Zhixiong Niu, Mingchao Zhang, Zhiying Xu, Chuntao Hu, Pengzhi Zhu, Qingchun Song, Peng Cheng 0005, Cam-Tu Nguyen, Shaoling Sun, Xiaohu Xu, Yongqiang Xiong, Wei Wang 0002, Xiaoliang Wang 0001 |
INFOCOM | 13 |
| 2025 | Software-based Live Migration for RDMAabstractLive migration is critical to ensure services are not interrupted during host maintenance in data centers. On the other hand, RDMA has been widely adopted in data centers, and has attracted both academia and industry for years. However, live migration of RDMA is not supported in today's data centers. Although modifying RDMA NICs (RNICs) to be aware of live migration has been proposed for years, there is no sign of supporting it on commodity RNICs. This paper proposes MigrRDMA, a software-based RDMA live migration that does not rely on any extra hardware support. MigrRDMA provides a software indirection layer to achieve transparent switching to new RDMA communications. Unlike previous RDMA virtualization that provides sharing and isolation, MigrRDMA's indirection layer focuses on keeping the RDMA states on the migration source and destination identical from the perspective of applications. We implemented MigrRDMA prototype over Mellanox RNICs. Our evaluation shows that MigrRDMA adds little downtime when migrating a container with live RDMA connections running at line rate. Besides, the MigrRDMA virtualization layer only adds 3% ~ 9% extra overheads in the data path. When migrating Hadoop tasks, MigrRDMA only incurs an extra 3-second job completion time. Ran Shu 0001, Yongqiang Xiong, Fengyuan Ren |
SIGCOMM | 3 |
| 2025 | HyperDrive: Direct Network Telemetry Storage via Programmable SwitchesabstractIn cloud datacenter operations, telemetry and logs are indispensable, enabling essential services such as network diagnostics, auditing, and knowledge discovery. The escalating scale of data centers, coupled with increased bandwidth and finer-grained telemetry, results in an overwhelming volume of data. This proliferation poses significant storage challenges for telemetry systems. In this article, we introduce HyperDrive, an innovative system designed to efficiently store large volumes of telemetry and logs in data centers using programmable switches. This in-network approach effectively mitigates bandwidth bottlenecks commonly associated with traditional endpoint-based methods. To our knowledge, we are the first to use a programmable switch to directly control storage, bypassing the CPU to achieve the best performance. With merely 21% of a switch’s resources, our HyperDrive implementation showcases remarkable scalability and efficiency. Through rigorous evaluation, it has demonstrated linear scaling capabilities, efficiently managing 12 SSDs on a single server with minimal host overhead. In an eight-server testbed, HyperDrive achieved an impressive throughput of approximately 730 Gbps, underscoring its potential to transform data center telemetry and logging practices. Ziyuan Liu 0008, Zhixiong Niu, Ran Shu 0001, Wenxue Cheng, Jacob Nelson 0001, Dan R. K. Ports, Peng Cheng 0005, Yongqiang Xiong |
IEEE Trans. Cloud Comput. | 9 |
| 2025 | Low-Overhead Intra-Host Container Communication With Hardware OffloadingabstractContainers are widely embraced for their deployment and performance benefits over virtual machines. Yet, for many data-intensive applications in containerized clouds, bulky data transfers may impose performance issues. In particular, communication across co-located containers on the same host incurs large overheads in memory copy and the kernel’s TCP stack. Existing solutions such as shared-memory networking and RDMA have their own limitations, including insufficient memory isolation and limited scalability. This paper presents PipeDevice, a new system for low overhead intra-host container communication. PipeDevice follows a hardware-software co-design approach — it offloads data forwarding entirely onto hardware, which accesses application data in hugepages on the host, thereby eliminating CPU overhead from memory copy and TCP processing. PipeDevice preserves memory isolation and scales well to connections, making it deployable in public clouds. Isolation is achieved by allocating dedicated memory to each connection from hugepages. To achieve high scalability, PipeDevice stores the connection states entirely in host DRAM and manages them in software. Evaluation with a prototype implementation on commodity FPGA shows that for delivering 80Gbps across containers PipeDevice saves 63.2% CPU compared to kernel TCP stack, and 40.5% over FreeFlow. PipeDevice provides salient benefits to applications. For example, we port baidu-allreduce to PipeDevice and obtain$\sim 2.2\times $gains in allreduce throughput. Zhixiong Niu, Ran Shu 0001, Peng Cheng 0005, Yongqiang Xiong, Dongsu Han, Chun Jason Xue, Hong Xu 0001 |
IEEE Trans. Netw. | 5 |
| 2024 | Software-based Live Migration for Containerized RDMAabstractContainer live migration is critical to ensure services are not interrupted during host maintenance in data centers. On the other hand, RDMA containerization has attracted both academia and industry for years. However, live migration for containerized RDMA is not supported in today’s data centers. Although modifying RDMA NICs (RNICs) to be aware of live migration has been proposed for years, there is no sign of supporting it on commodity RNICs. This paper proposes MigrRDMA, a software-based RDMA live migration for containers, which does not rely on any extra hardware support. MigrRDMA provides a minimum virtualization layer inside the RDMA library loaded in applications, which achieves transparent switching to new RDMA communications. Unlike previous RDMA virtualization that provides sharing and isolation, MigrRDMA’s virtualization layer focuses on keeping the RDMA states on the migration source and destination the same from the perspective of applications. Our evaluation shows that MigrRDMA only adds 0.7 ∼ 12.1 ms downtime to migrate a container with live RDMA connections running at line rate. Besides, the MigrRDMA virtualization layer only adds 3% ∼ 9% overheads in the data path operations. Ran Shu 0001, Yongqiang Xiong, Fengyuan Ren |
APNet | 3 |
| 2024 | Bridging Gaps in LLM Code Translation: Reducing Errors with Call Graphs and Bridged DebuggersabstractWhen using large language models (LLMs) for code translation of complex software, numerous compilation and runtime errors can occur due to insufficient context awareness. To address this issue, this paper presents a code translation method based on call graphs and bridged debuggers: TransGraph. TransGraph first obtains the call graph of the entire code project using the Language Server Protocol, which provides a detailed description of the function call relationships in the program. Through this structured view of the code, LLMs can more effectively handle large-scale and complex codebases, significantly reducing compilation errors. Furthermore, TransGraph, combined with bridged debuggers and dynamic test case generation, significantly reduces runtime errors, overcoming the limitations of insufficient test case coverage in traditional methods. In experiments on six datasets including CodeNet and Avatar, TransGraph outperformed existing code translation methods and LLMs in terms of translation accuracy, with improvements of up to 10.2%. Fajun Zhang, Ling Liang 0002, Yongqiang Xiong |
ASE | 5 |
| 2024 | NeoMem: Hardware/Software Co-Design for CXL-Native Memory TieringabstractThe Compute Express Link (CXL) interconnect makes it feasible to integrate diverse types of memory into servers via its byte-addressable SerDes links. Considering the various access latency, harnessing the full potential of CXL-based heterogeneous memory systems requires efficient memory tiering. However, prior work can hardly make a fundamental progress owing to low-resolution and high-overhead memory access profiling techniques. To address this critical challenge, we propose a novel memory tiering solution called NeoMem, which features a hardware/software co-design. NeoMem offloads memory profiling functions to CXL device-side controllers, integrating a dedicated hardware unit called NeoProf. NeoProf readily monitors memory accesses and provides the OS with crucial page hotness statistics and other useful system state information. On the OS kernel side, we design a revamped memory-tiering strategy, enabling accurate and timely hot page promotion based on NeoProf statistics. We implement NeoMem on a real FPGA-based CXL memory platform and Linux kernel v6.3. Comprehensive evaluations demonstrate that NeoMem achieves 32% ~ 67% geomean speedup over several existing memory tiering solutions. Zhe Zhou 0002, Tao Zhang 0032, Yang Wang 0053, Ran Shu 0001, Shuotao Xu, Peng Cheng 0005, Yongqiang Xiong, Jie Zhang 0048, Guangyu Sun 0003 |
MICRO | 9 |
| 2024 | SuperBench: Improving Cloud AI Infrastructure Reliability with Proactive Validation
Yifan Xiong 0001, Ziyue Yang 0002, Guoshuai Zhao 0001, Dong Zhong, Boris Pinzur, Jie Zhang 0048, Yang Wang 0053, Hossein Pourreza, Jeff Baxter, Kushal Datta, Prabhat Ram, Luke Melton, Joe Chau, Peng Cheng 0005, Yongqiang Xiong, Lidong Zhou |
USENIX ATC | 19 |
| 2024 | Intelligent Packet Processing for Performant Containers in IoTabstractThis article explores the computing and communication overhead of network processing in Internet of Things (IoT) devices, focusing on containers, a major building block for the edge computing. Our experiments reveal that containers on IoT devices suffer$\sim 2.6\times $higher CPU usage for SoftIRQ processing, ~59% less network throughput, and$2\times $higher per-packet latency on average than native processes. While several existing studies enhance networking performance, they often sacrifice interoperability by requiring special hardware or modifying networking semantics or APIs. Thus, we design and implement a kernel networking accelerator, called SCON, that maintains interoperability, crucial for IoT devices. SCON addresses major bottlenecks in container networking through system-level profiling. We evaluate SCON with three types of IoT devices. On the Raspberry Pi 4, SCON reduces the latencies of major IoT application protocols (e.g., HTTP and MQTT) by$\sim 10\times $, achieving a similar level of latency to the native process. Further analysis shows that SCON reduces CPU usage for SoftIRQ processing by ~26%. We also report similar improvements on the other two IoT devices. Our conclusion is that SCON is unique in significantly reducing the computing and communication overhead of container networking in IoT devices while maintaining interoperability. Furthermore, it works consistently across different types of devices, whether wired or wireless, and regardless of heavy or sporadic traffic. Wonmi Choi, Yeonho Yoo, Kyungwoon Lee, Zhixiong Niu, Peng Cheng 0005, Yongqiang Xiong, Gyeongsik Yang, Chuck Yoo |
IEEE Internet Things J. | 6 |
| 2023 | MINA: Auto-scale In-network Aggregation for Machine Learning Service
Shichen Dong, Zhixiong Niu, Mingchao Zhang, Zhiying Xu, Chuntao Hu, Wei Wang 0002, Pengzhi Zhu, Qingchun Song, Peng Cheng 0005, Yongqiang Xiong, Chen Tian 0001, Cam-Tu Nguyen, Xiaoliang Wang 0001 |
APNet | 11 |
| 2023 | SlimeMold: Hardware Load Balancer at Scale in DatacenterabstractStateful load balancers (LB) are essential services in cloud data centers, playing a crucial role in enhancing the availability and capacity of applications. Numerous studies have proposed methods to improve the throughput, connections per second, and concurrent flows of single LBs. For instance, with the advancement of programmable switches, hardware-based load balancers (HLB) have become mainstream due to their high efficiency. However, programmable switches still face the issue of limited registers and table entries, preventing them from fully meeting the performance requirements of data centers. In this paper, rather than solely focusing on enhancing individual HLBs, we introduce SlimeMold, which enables HLBs to work collaboratively at scale as an integrated LB system in data centers. Ziyuan Liu 0008, Zhixiong Niu, Ran Shu 0001, Guohong Lai, Zongying He, Jacob Nelson 0001, Dan R. K. Ports, Peng Cheng 0005, Yongqiang Xiong |
APNet | 12 |
| 2023 | SegaNet: An Advanced IoT Cloud Gateway for Performant and Priority-Oriented Message DeliveryabstractWith the tremendous growth of IoT, the role of IoT cloud gateways in facilitating communication between IoT devices and the cloud has become more important than ever before. Most previous studies have focused on developing interoperability between IoT and cloud to accommodate various radio protocols. However, they have often neglected the performance aspect of the IoT cloud gateway, leaving users with limited options: either purchasing multiple gateways or connecting only a small number of IoT devices. Through our comprehensive measurements and analysis, we identified five key issues in IoT cloud gateways related to high latency, CPU bottlenecks, inefficient network stacks on ARM, substantial encryption overhead, and the lack of priority support. To address these issues, we propose a new IoT cloud gateway - SegaNet. We carefully design with 1) multiple agents management, 2) efficient TLS encryption, and 3) priority-oriented message delivery. Our prototype evaluation shows up to 16.7 × lower latency and 4.5 × lower CPU consumption than gateways of the existing IoT-cloud ecosystem. Yeonho Yoo, Zhixiong Niu, Chuck Yoo, Peng Cheng 0005, Yongqiang Xiong |
APNet | 5 |
| 2023 | Polaris: Enhancing CXL-based Memory Expanders with Memory-side Prefetching
Zhe Zhou 0002, Shuotao Xu, Tao Zhang 0032, Ran Shu 0001, Peng Cheng 0005, Yongqiang Xiong, Guangyu Sun 0003 |
APPT | 8 |
| 2023 | ARK: GPU-driven Code Execution for Distributed Deep Learning
Changho Hwang, KyoungSoo Park, Ran Shu 0001, Xinyuan Qu, Peng Cheng 0005, Yongqiang Xiong |
NSDI | 6 |
| 2023 | Poster: Meili: Towards SmartNIC as a ServiceabstractThe gap between the stagnation of CPU power and the increase in network bandwidth has promoted a shift towards placing more computation on network hardware [16, 17]. Therefore, SmartNICs have become prevalent in data centers to serve various cloud applications, from network functions [15, 17, 22] to high-level applications like distributed applications and storage [14, 16, 18--21, 23]. Shaofeng Wu, Zhixiong Niu, Ran Shu 0001, Peng Cheng 0005, Yongqiang Xiong, Chun Jason Xue, Zaoxing Liu, Hong Xu 0001 |
SIGCOMM | 6 |
| 2022 | OpenNetLab: Open Platform for RL-based Congestion Control for Real-Time CommunicationsabstractWith the growing importance of real-time communications (RTC), designing congestion control (CC) algorithms for RTC that achieve high network performance and QoE is gaining attention. Recently, data-driven, reinforcement learning (RL)-based CC algorithms for RTC have shown great potential, outperforming traditional rule-based counterparts. However, there are no open platforms tailored for training, evaluation, and validation of the algorithms that can facilitate this emerging research area. Jeongyoon Eo, Zhixiong Niu, Wenxue Cheng, Francis Y. Yan, Jorina Kardhashi, Scott Inglis, Michael Revow, Byung-Gon Chun, Peng Cheng 0005, Yongqiang Xiong |
APNet | 11 |
| 2022 | A Disaggregate Data Collecting Approach for Loss-Tolerant ApplicationsabstractDatacenter generates operation data at an extremely high rate, and data center operators collect and analyze them for problem diagnosis, resource utilization improvement, and performance optimization. However, existing data collection methods fail to efficiently aggregate and store data at extremely high speed and scale. In this paper, we explore a new approach that leverages programmable switches to aggregate data and directly write data to the destination storage. Our proposed data collection system, ALT, uses programmable switches to control NVMe SSDs on remote hosts without the involvement of a remote CPU. To tolerate loss, ALT uses an elegant data structure to enable efficient data recovery when retrieving the collected data. We implement our system on a Tofino-based programmable switch for a prototype. Our evaluation shows that ALT can saturate SSD’s peak performance without any CPU involvement. Ziyuan Liu 0008, Zhixiong Niu, Ran Shu 0001, Wenxue Cheng, Peng Cheng 0005, Yongqiang Xiong, Jacob Nelson 0001, Dan R. K. Ports |
APNet | 6 |
| 2022 | PipeDevice: a hardware-software co-design approach to intra-host container communicationabstractContainers are prevalently adopted due to the deployment and performance advantages over virtual machines. For many containerized data-intensive applications, however, bulky data transfers may pose performance issues. In particular, communication across co-located containers on the same host incurs large overheads in memory copy and the kernel's TCP stack. Existing solutions such as shared-memory networking and RDMA have their own limitations, including insufficient memory isolation and limited scalability. Chuanwen Wang, Zhixiong Niu, Ran Shu 0001, Peng Cheng 0005, Yongqiang Xiong, Dongsu Han, Chun Jason Xue, Hong Xu 0001 |
CoNEXT | 6 |
| 2022 | Moneo: Non-intrusive Fine-grained Monitor for AI InfrastructureabstractCloud-based AI infrastructure is increasingly important, especially on large-scale distributed training. To improve its efficiency and serviceability, real-time monitoring of the infrastructure and profiling the workload are proved to be the effective approach empirically. However, cloud environment poses great challenges as service providers cannot interfere with their tenants' workloads or touch user data, thus previous instrumentation-based monitoring approach cannot be applied, nor does the workload trace collection.We propose Moneo, a non-intrusive cloud-friendly monitoring system for AI infrastructure. Moneo is capable of intelligently collecting the key architecture-level metrics at finer granularity in real-time without instrumenting or tracing the workloads, which has been deployed in real production cloud, Azure. We analyze the results reported by Moneo for typical large-scale distributed AI workloads from real deployment. Results demonstrate that Moneo can effectively help service providers understand the real resource usage patterns of various AI workloads and real networking requirements, so as to get valuable findings help improve the efficiency of cloud infrastructure and optimize the software stack with the consideration of the characteristic resource usage requirements for different AI workloads. Yifan Xiong 0001, Chen Tian 0001, Peng Cheng 0005, Yongqiang Xiong |
ICC | 7 |
| 2022 | RuleCache: Accelerating Web Application Firewalls by On-line Learning Traffic PatternsabstractWeb Application Firewall (WAF) is widely deployed in cloud to protect web applications, whose performance becomes one of the major bottlenecks for web services. In this paper, we comprehensively analyze several root causes that downgrade WAF’s efficiency. Inspired by that, we build a caching system RuleCache to devise optimization strategies for improving WAF’s performance. Among, Rule Ordering Cache is online learning an optimal order of the ruleset for a better performance of blocking. Rule Result Cache reuses rule results of targets, saving large repetitive computations. Additionally, Rule Prepruning Cache aims to cut extra overhead by processing the static rules in the offline stage. Our evaluation demonstrates that the prototype can improve the performance by up to 3.85x, 1.57x, and 2.4x respectively with the above modules, and up to 5.5x in total. Qingni Shen, Peng Cheng 0005, Yongqiang Xiong, Zhonghai Wu |
ICWS | 4 |
| 2022 | An Adaptive Deep RL Method for Non-Stationary Environments with Piecewise Stable ContextabstractOne of the key challenges in deploying RL to real-world applications is to adapt to variations of unknown environment contexts, such as changing terrains in robotic tasks and fluctuated bandwidth in congestion control. Existing works on adaptation to unknown environment contexts either assume the contexts are the same for the whole episode or assume the context variables are Markovian. However, in many real-world applications, the environment context usually stays stable for a stochastic period and then changes in an abrupt and unpredictable manner within an episode, resulting in a segment structure, which existing works fail to address. To leverage the segment structure of piecewise stable context in real-world applications, in this paper, we propose a \textit{\textbf{Se}gmented \textbf{C}ontext \textbf{B}elief \textbf{A}ugmented \textbf{D}eep~(SeCBAD)} RL method. Our method can jointly infer the belief distribution over latent context with the posterior over segment length and perform more accurate belief context inference with observed data within the current context segment. The inferred belief context can be leveraged to augment the state, leading to a policy that can adapt to abrupt variations in context. We demonstrate empirically that SeCBAD can infer context segment length accurately and outperform existing methods on a toy grid world environment and Mujuco tasks with piecewise-stable context. Xiangming Zhu 0002, Pushi Zhang, Li Zhao 0007, Wenxue Cheng, Peng Cheng 0005, Yongqiang Xiong, Tao Qin 0001, Jianyu Chen 0002, Tie-Yan Liu |
NeurIPS | 8 |
| 2022 | NetKernel: Making Network Stack Part of the Virtualized InfrastructureabstractThis paper presents a system called NetKernel that decouples the network stack from the guest virtual machine and offers it as an independent module. NetKernel represents a new paradigm where network stack can be managed as part of the virtualized infrastructure. It provides important efficiency benefits: By gaining control and visibility of the network stack, operators can perform network management more directly and flexibly, such as multiplexing VMs running different applications to the same network stack module to save CPU cores, and enforcing fair bandwidth sharing. Users also benefit from the simplified stack deployment and better performance: For example mTCP can be deployed without API change to support nginx natively, and shared memory networking can be readily enabled to improve performance of colocated VMs. Testbed evaluation using 100G NICs shows that NetKernel preserves the performance and scalability of both kernel and userspace network stacks, and provides the same isolation as the current architecture. Zhixiong Niu, Peng Cheng 0005, Yongqiang Xiong, Dongsu Han, Keith Winstein, Chun Jason Xue, Hong Xu 0001 |
IEEE/ACM Trans. Netw. | 4 |
| 2022 | Stanza: Layer Separation for Distributed Training in Deep LearningabstractThe parameter server architecture is prevalently used for distributed deep learning. Each worker machine in a such system trains the complete model, which leads to a large amount of network data transfer between workers and servers. We empirically observe that the data transfer has a major impact on training time. We present a new distributed training system called Stanza to tackle this problem. Stanza exploits the fact that in many models such as convolution neural networks, most data exchange is attributed to the fully connected layers, while most computation is carried out in convolutional layers. Thus, we propose layer separation in distributed training: most nodes of the cluster train only the convolutional layers, while the rest train the fully connected layers. Gradients and parameters of the fully connected layers no longer need to be exchanged across the entire cluster, thereby substantially reducing the data transfer volume. We implement Stanza on PyTorch and evaluate its performance on Azure and EC2. Results show that Stanza accelerates training significantly over current parameter server systems: on EC2 instances with Tesla V100 GPU and 10Gb bandwidth for example, Stanza is 1.34x–13.9x faster for common deep learning models. Xiaorui Wu, Hong Xu 0001, Bo Li 0001, Yongqiang Xiong |
IEEE Trans. Serv. Comput. | 4 |
| 2021 | NFD: Using Behavior Models to Develop Cross-Platform Network FunctionsabstractNFV ecosystem is flourishing and more and more NF platforms appear, but this makes NF vendors difficult to deliver NFs rapidly to diverse platforms. We propose an NF development framework named NFD for cross-platform NF development. NFD's main idea is to decouple the functional logic from the platform logic -it provides a platform-independent language to program NFs' behavior models, and a compiler with interfaces to develop platform-specific plugins. By enabling a plugin on the compiler, various NF models would be compiled to executables integrated with the target platform. We prototype NFD, build 14 NFs, and support 6 platforms (standard Linux, OpenNetVM, GPU, SGX, DPDK, OpenNF). Our evaluation shows that NFD can save development workload for cross-platform NFs and output valid and performant NFs. Hongyi Huang, Wenfei Wu, Yongchao He, Bangwen Deng, Ying Zhang 0022, Yongqiang Xiong, Guo Chen 0001, Yong Cui 0001, Peng Cheng 0005 |
INFOCOM | 6 |
| 2021 | One More Config is Enough: Saving (DC)TCP for High-Speed Extremely Shallow-Buffered DatacentersabstractThe link speed in production datacenters is growing fast, from 1 Gbps to 40 Gbps or even 100 Gbps. However, the buffer size of commodity switches increases slowly, e.g., from 4 MB at 1 Gbps to 16 MB at 100 Gbps, thus significantly outpaced by the link speed. In such extremely shallow-buffered networks, today's TCP/ECN solutions, such as DCTCP, suffer from either excessive packet losses or significant throughput degradation. Motivated by this, we introduce BCC,1a simple yet effective solution that requires only one more ECN configuration (i.e., shared buffer ECN/RED) at commodity switches. BCC operates upon real-time global shared buffer utilization. When available buffer space suffices, BCC delivers both high throughput and low packet loss rate as prior work; When it gets insufficient, BCC automatically triggers the shared buffer ECN to prevent packet loss at the cost of sacrificing a small amount of throughput. BCC is readily deployable with existing commodity switches. We validate BCC's efficacy in a 100G testbed and evaluate its performance using extensive simulations. Our results show that BCC maintains low packet loss rate persistently while only slightly degrading throughput when the buffer becomes insufficient. For example, compared to current practice, BCC achieves up to 94.4% lower 99th percentile flow completion time (FCT) for small flows while only degrading average FCT for large flows by up to 3%. Wei Bai 0001, Shuihai Hu, Kai Chen 0005, Kun Tan 0002, Yongqiang Xiong |
IEEE/ACM Trans. Netw. | 5 |
| 2021 | Accelerating End-to-End Deep Learning Workflow With Codesign of Data Preprocessing and SchedulingabstractIn this article, we investigate the performance bottleneck of existing deep learning (DL) systems and propose DLBooster to improve the running efficiency of deploying DL applications on GPU clusters. At its core, DLBooster leverages two-level optimizations to boost the end-to-end DL workflow. On the one hand, DLBooster selectively offloads some key decoding workloads to FPGAs to provide high-performance online data preprocessing services to the computing engine. On the other hand, DLBooster reorganizes the computational workloads of training neural networks with the backpropagation algorithm and schedules them according to their dependencies to improve the utilization of GPUs at runtime. Based on our experiments, we demonstrate that compared with baselines, DLBooster can improve the image processing throughput by 1.4× - 2.5× and reduce the processing latency by 1/3 in several real-world DL applications and datasets. Moreover, DLBooster consumes less than 1 CPU core to manage FPGA devices at runtime, which is at least 90 percent less than the baselines in some cases. DLBooster shows its potential to accelerate DL workflows in the cloud. Dan Li 0001, Binyao Jiang, Jinkun Geng, Wei Bai 0001, Yongqiang Xiong |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2020 | One More Config is Enough: Saving (DC)TCP for High-speed Extremely Shallow-buffered DatacentersabstractThe link speed in production datacenters is growing fast, from 1Gbps to 40Gbps or even 100Gbps. However, the buffer size of commodity switches increases slowly, e.g., from 4MB at 1Gbps to 16MB at 100Gbps, thus significantly outpaced by the link speed. In such extremely shallow-buffered networks, today's TCP/ECN solutions, such as DCTCP, suffer from either excessive packet loss or substantial throughput degradation.To this end, we present BCC1, a simple yet effective solution that requires just one more ECN config (i.e., shared buffer ECN/RED) over prior solutions. BCC operates based on real-time global shared buffer utilization. When available buffer space suffices, BCC delivers both high throughput and low packet loss rate as prior work; Once it gets insufficient, BCC automatically triggers the shared buffer ECN to prevent packet loss at the cost of sacrificing little throughput. BCC is readily deployable with existing commodity switches. We validate BCC's hardware feasibility in a small 100G testbed and evaluate its performance using large-scale simulations. Our results show that BCC maintains low packet loss rate while slightly degrading throughput when the available buffer becomes insufficient. For example, compared to current practice, BCC achieves up to 94.4% lower 99th percentile flow completion time (FCT) for small flows while degrading average FCT for large flows by up to 3%. Wei Bai 0001, Shuihai Hu, Kai Chen 0005, Kun Tan 0002, Yongqiang Xiong |
INFOCOM | 5 |
| 2020 | NetKernel: Making Network Stack Part of the Virtualized Infrastructure
Zhixiong Niu, Hong Xu 0001, Peng Cheng 0005, Yongqiang Xiong, Tao Wang 0088, Dongsu Han, Keith Winstein |
USENIX ATC | 5 |
| 2019 | URSA: Hybrid Block Storage for Cloud-Scale Virtual DisksabstractThis paper presents URSA, a hybrid block store that provides virtual disks for various applications to run efficiently on cloud VMs. Trace analysis shows that the I/O patterns served by block storage have limited locality to exploit. Therefore, instead of using SSDs as a cache layer, URSA proposes an SSD-HDD-hybrid storage structure that directly stores primary replicas on SSDs and replicates backup replicas on HDDs, using journals to bridge the performance gap between SSDs and HDDs. URSA integrates the hybrid structure with designs for high reliability, scalability, and availability. Experiments show that URSA in its hybrid mode achieves almost the same performance as in its SSD-only mode (storing all replicas on SSDs), and outperforms other block stores (Ceph and Sheepdog) even in their SSD-only mode while achieving much higher CPU efficiency (performance per core). We also discuss some practical issues in our deployment. Huiba Li, Yiming Zhang 0003, Dongsheng Li 0001, Shengyun Liu, Peng Huang 0005, Zheng Qin 0002, Kai Chen 0005, Yongqiang Xiong |
EuroSys | 9 |
| 2019 | FlowShader: a Generalized Framework for GPU-accelerated VNF Flow ProcessingabstractGPU acceleration has been widely investigated for packet processing in virtual network functions (NFs), but not for L7 flow-processing NFs. In L7 NFs, reassembled TCP messages of the same flow should be processed in order in the same processing thread, and the uneven sizes among flows pose a major challenge for full realization of GPU's parallel computation power. To exploit GPUs for L7 NF processing, this paper presents FlowShader, a GPU acceleration framework to achieve both high generality and throughput even under skewed flow size distributions. We carefully design an efficient scheduling algorithm that fully exploits available GPU and CPU capacities; in particular, we dispatch large flows which seriously break up the size balance to CPU and the rest of flows to GPU. Furthermore, FlowShader allows similar NF logic (as CPU-based NFs) to run on individual threads in a GPU, which is more generalized and easy to take on as compared to redesigning an NF for operation parallelism on GPU. We implemented a number of L7 flow processing NFs based on FlowShader. Evaluations are conducted under both synthetic and real-world traffic traces and results show that the throughput achieved by FlowShader is up to 6x that of the CPU-only baseline and 3x of the GPU-only design. Xiaodong Yi 0001, Jingpu Duan, Wei Bai 0001, Chuan Wu 0001, Yongqiang Xiong, Dongsu Han |
ICNP | 6 |
| 2019 | DLBooster: Boosting End-to-End Deep Learning Workflows with Offloading Data Preprocessing PipelinesabstractIn recent years, deep learning (DL) has prospered again due to improvements in both computing and learning theory. Emerging studies mostly focus on the acceleration of refining DL models but ignore data preprocessing issues. However, data preprocessing can significantly affect the overall performance of end-to-end DL workflows. Our studies on several image DL workloads show that existing preprocessing backends are quite inefficient: they either perform poorly in throughput (30% degradation) or burn too many (>10) CPU cores. Based on these observations, we propose DLBooster, a high-performance data preprocessing pipeline that selectively offloads key workloads to FPGAs, to fit the stringent demands on data preprocessing for cutting-edge DL applications. Our testbed experiments show that, compared with the existing baselines, DLBooster can achieve 1.35×~2.4× image processing throughput in several DL workloads, but consumes only 1/10 CPU cores. Besides, it also reduces the latency by 1/3 in online image inference. Dan Li 0001, Binyao Jiang, Xi Fan, Jinkun Geng, Wei Bai 0001, Ran Shu 0001, Peng Cheng 0005, Yongqiang Xiong |
ICPP | 13 |
| 2019 | Direct Universal Access: Making Data Center Resources Available to FPGA
Ran Shu 0001, Peng Cheng 0005, Guo Chen 0001, Yongqiang Xiong, Derek Chiou, Thomas Moscibroda |
NSDI | 6 |
| 2019 | Accelerating Rule-matching Systems with Learned Rankers
Zhao Lucis Li, Chieh-Jan Mike Liang, Wei Bai 0001, Yongqiang Xiong, Guangzhong Sun |
USENIX ATC | 5 |
| 2019 | MP-RDMA: Enabling RDMA With Multi-Path Transport in DatacentersabstractRDMA is becoming prevalent because of its low latency, high throughput and low CPU overhead. However, in current datacenters, RDMA remains a single path transport which is prone to failures and falls short to utilize the rich parallel network paths. Unlike previous multi-path approaches, which mainly focus on TCP, this paper presents a multi-path transport for RDMA, i.e. MP-RDMA, which efficiently utilizes the rich network paths in datacenters. MP-RDMA employs three novel techniques to address the challenge of limited RDMA NICs on-chip memory size: 1) a multi-path ACK-clocking mechanism to distribute traffic in a congestion-aware manner without incurring per-path states; 2) an out-of-order aware path selection mechanism to control the level of out-of-order delivered packets, thus minimizes the meta data required to them; 3) a synchronise mechanism to ensure in-order memory update whenever needed. With all these techniques, MP-RDMA only adds 66B to each connection state compared to single-path RDMA. Our evaluation with an FPGA-based prototype demonstrates that compared with single-path RDMA, MP-RDMA can significantly improve the robustness under failures ( $2\times \sim 4\times $ higher throughput under 0.5%~10% link loss ratio) and improve the overall network utilization by up to 47%. Guo Chen 0001, Yuanwei Lu, Bojie Li, Kun Tan 0002, Yongqiang Xiong, Peng Cheng 0005, Jiansong Zhang 0001, Thomas Moscibroda |
IEEE/ACM Trans. Netw. | 5 |
| 2018 | The Fusion of VMs and Processes: A System Perspective of cKernelabstractVirtual machines (VMs) and processes are two important abstractions for cloud virtualization, where VMs usually install a complete operating system (OS) executing user processes. Although existing in different layers in the virtualization hierarchy, VMs and processes have overlapped functionalities. For example, they are both intended to provide execution abstraction (e.g., physical/virtual memory address space), and share similar objectives of isolation, cooperation and scheduling. However, neither of them could provide the benefits of the other: VMs provide higher isolation, security and portability, while processes are more efficient, flexible and easier to schedule and cooperate. Currently, this heavyweight architecture degrades both efficiency and security of cloud services. There are two trends for cloud virtualization: the first is to enhance processes to achieve VM-like security, and the second is to reduce VMs to achieve process-like flexibility. Based on these observations, our vision is that in the near future VMs and processes might be fused into one new abstraction for cloud virtualization that embraces the best of both, providing VM-level isolation and security while preserving process-level efficiency and flexibility. We describe a reference implementation, dubbed cKernel (customized kernel), for the new abstraction. Essentially, cKernel enhances the exokernel architecture by (i) adopting the LibOS paradigm to assemble isolated, smallest possible "execution environments", and (ii) following the the "core-shell" model to dynamically add traditional process features to the environments. Yiming Zhang 0003, Dongsheng Li 0001, Yingwen Chen 0001, Ping Zhong 0002, Yongqiang Xiong, Huaimin Wang 0001 |
ICDCS | 8 |
| 2018 | Multi-Path Transport for RDMA in Datacenters
Yuanwei Lu, Guo Chen 0001, Bojie Li, Kun Tan 0002, Yongqiang Xiong, Peng Cheng 0005, Jiansong Zhang 0001, Enhong Chen, Thomas Moscibroda |
NSDI | 5 |
| 2018 | KylinX: A Dynamic Library Operating System for Simplified and Efficient Cloud Virtualization
Yiming Zhang 0003, Jon Crowcroft, Dongsheng Li 0001, Chengfen Zhang, Huiba Li, Yaozheng Wang, Yongqiang Xiong, Guihai Chen |
USENIX ATC | 8 |
| 2018 | FUSO: Fast Multi-Path Loss Recovery for Data Center Networks
Guo Chen 0001, Yuanwei Lu, Yuan Meng 0002, Bojie Li, Kun Tan 0002, Dan Pei, Peng Cheng 0005, Layong Luo, Yongqiang Xiong, Xiaoliang Wang 0001, Youjian Zhao |
IEEE/ACM Trans. Netw. | 9 |
| 2017 | Congestion Control for High-speed Extremely Shallow-buffered Datacenter NetworksabstractThe link speed in datacenters is growing fast, from 1Gbps to 100Gbps. However, the buffer size of commodity switches increases slowly, thus significantly outpaced by the link speed. In such extremely shallow-buffered datacenter networks, prior TCP/ECN solutions suffer from either excessive packet losses or significant throughput degradation. Motivated by this, we introduce BCC, a simple yet effective solution with only one more configuration (shared buffer ECN/RED) at commodity switches. BCC operates based on real-time shared buffer utilization. When the buffer is abundant, BCC delivers both high throughput and low packet loss rate. When it becomes scarce, BCC triggers shared buffer ECN/RED to prevent packet losses at the cost of sacrificing a small amount of throughput. Our preliminary results show that BCC maintains low packet loss rate persistently while only slightly degrading throughput when the buffer becomes insufficient. Compared to current practice, BCC achieves up to 94.4% lower 99th percentile completion time for small flows while only degrading large flows by up to 2.8%. Wei Bai 0001, Kai Chen 0005, Shuihai Hu, Kun Tan 0002, Yongqiang Xiong |
APNet | 5 |
| 2017 | Memory Efficient Loss Recovery for Hardware-based Transport in DatacenterabstractLimited by the small on-chip memory, hardware-based transport typically implements go-back-N loss recovery mechanism, which costs very few memory but is well-known to perform inferior even under small packet loss ratio. We present MELO, an efficient selective retransmission mechanism for hardware-based transport, which consumes only a constant small memory regardless of the number of concurrent connections. Specifically, MELO employs an architectural separation between data and meta data storage and uses a shared bits pool allocation mechanism to reduce meta data on-chip memory footprint. By only adding in average 23B extra on-chip states for each connection, MELO achieves up to 14.02x throughput while reduces 99% tail FCT by 3.11x compared with go-back-N under certain loss ratio. Yuanwei Lu, Guo Chen 0001, Zhenyuan Ruan, Wencong Xiao, Bojie Li, Jiansong Zhang 0001, Yongqiang Xiong, Peng Cheng 0005, Enhong Chen |
APNet | 7 |
| 2017 | Network Stack as a Service in the CloudabstractThe tenant network stack is implemented inside the virtual machines in today's public cloud. This legacy architecture presents a barrier to protocol stack innovation due to the tight coupling between the network stack and the guest OS. In particular, it causes many deployment troubles to tenants and management and efficiency problems to the cloud provider. To address these issues, we articulate a vision of providing the network stack as a service. The central idea is to decouple the network stack from the guest OS, and offer it as an independent entity implemented by the cloud provider. This re-architecting allows tenants to readily deploy any stack independent of its kernel, and the provider to offer meaningful SLAs to tenants by gaining control over the network stack. We sketch an initial design called NetKernel to accomplish this vision. Our preliminary testbed evaluation with a prototype shows the feasibility and benefits of our idea. Zhixiong Niu, Hong Xu 0001, Dongsu Han, Peng Cheng 0005, Yongqiang Xiong, Guo Chen 0001, Keith Winstein |
HotNets | 5 |
| 2017 | One more queue is enough: Minimizing flow completion time with explicit priority notificationabstractIdeally, minimizing the flow completion time (FCT) requires millions of priorities supported by the underlying network so that each flow has its unique priority. However, in production datacenters, the available switch priority queues for flow scheduling are very limited (merely 2 or 3). This practical constraint seriously degrades the performance of previous approaches. In this paper, we introduce Explicit Priority Notification (EPN), a novel scheduling mechanism which emulates fine-grained priorities (i.e., desired priorities or DP) using only two switch priority queues. EPN can support various flow scheduling disciplines with or without flow size information. We have implemented EPN on commodity switches and evaluated its performance with both testbed experiments and extensive simulations. Our results show that, with flow size information, EPN achieves comparable FCT as pFabric that requires clean-slate switch hardware. And EPN also outperforms TCP by up to 60.5% if it bins the traffic into two priority queues according to flow size. In information-agnostic setting, EPN outperforms PIAS with two priority queues by up to 37.7%. To the best of our knowledge, EPN is the first system that provides millions of priorities for flow scheduling with commodity switches. Yuanwei Lu, Guo Chen 0001, Larry Luo, Kun Tan 0002, Yongqiang Xiong, Xiaoliang Wang 0001, Enhong Chen |
INFOCOM | 5 |
| 2017 | KV-Direct: High-Performance In-Memory Key-Value Store with Programmable NICabstractPerformance of in-memory key-value store (KVS) continues to be of great importance as modern KVS goes beyond the traditional object-caching workload and becomes a key infrastructure to support distributed main-memory computation in data centers. Recent years have witnessed a rapid increase of network bandwidth in data centers, shifting the bottleneck of most KVS from the network to the CPU. RDMA-capable NIC partly alleviates the problem, but the primitives provided by RDMA abstraction are rather limited. Meanwhile, programmable NICs become available in data centers, enabling in-network processing. In this paper, we present KV-Direct, a high performance KVS that leverages programmable NIC to extend RDMA primitives and enable remote direct key-value access to the main host memory. Bojie Li, Zhenyuan Ruan, Wencong Xiao, Yuanwei Lu, Yongqiang Xiong, Andrew Putnam, Enhong Chen |
SOSP | 5 |
| 2017 | Protego: Cloud-Scale Multitenant IPsec Gateway
Jeongseok Son, Yongqiang Xiong, Kun Tan 0002, Ze Gan, Sue B. Moon |
USENIX ATC | 2 |
| 2017 | Expeditus: Congestion-Aware Load Balancing in Clos Data Center NetworksabstractData center networks often use multi-rooted Clos topologies to provide a large number of equal cost paths between two hosts. Thus, load balancing traffic among the paths is important for high performance and low latency. However, it is well known that ECMP-the de facto load balancing scheme-performs poorly in data center networks. The main culprit of ECMP's problems is its congestion agnostic nature, which fundamentally limits its ability to deal with network dynamics. We propose Expeditus, a novel distributed congestion-aware load balancing protocol for general 3-tier Clos networks. The complex 3-tier Clos topologies present significant scalability challenges that make a simple per-path feedback approach infeasible. Expeditus addresses the challenges by using simple local information collection, where a switch only monitors its egress and ingress link loads. It further employs a novel two-stage path selection mechanism to aggregate relevant information across switches and make path selection decisions. Testbed evaluation on Emulab and large-scale ns-3 simulations demonstrate that, Expeditus outperforms ECMP by up to 45% in tail flow completion times (FCT) for mice flows, and by up to 38% in mean FCT for elephant flows in 3-tier Clos networks. Peng Wang 0037, Hong Xu 0001, Zhixiong Niu, Dongsu Han, Yongqiang Xiong |
IEEE/ACM Trans. Netw. | 5 |
| 2017 | CubicRing: Exploiting Network Proximity for Distributed In-Memory Key-Value StoreabstractIn-memory storage has the benefits of low I/O latency and high I/O throughput. Fast failure recovery is crucial for large-scale in-memory storage systems, bringing network-related challenges, including false detection due to transient network problems, traffic congestion during the recovery, and top-of-rack switch failures. In order to achieve fast failure recovery, in this paper, we present CubicRing, a distributed structure for cube-based networks, which exploits network proximity to restrict failure detection and recovery within the smallest possible one-hop range. We leverage the CubicRing structure to address the aforementioned challenges and design a network-aware in-memory key-value store called MemCube. In a 64-node 10GbE testbed, MemCube recovers 48 GB of data for a single server failure in 3.1 s. The 14 recovery servers achieve 123.9 Gb/s aggregate recovery throughput, which is 88.5% of the ideal aggregate bandwidth and several times faster than RAMCloud with the same configurations. Yiming Zhang 0003, Dongsheng Li 0001, Chuanxiong Guo, Yongqiang Xiong, Xicheng Lu |
IEEE/ACM Trans. Netw. | 5 |
| 2016 | Expeditus: Congestion-aware Load Balancing in Clos Data Center NetworksabstractData center networks often use multi-rooted Clos topologies to provide a large number of equal cost paths between two hosts. Thus, load balancing traffic among the paths is important for high performance and low latency. However, it is well known that ECMP---the de facto load balancing scheme---performs poorly in data center networks. The main culprit of ECMP's problems is its congestion agnostic nature, which fundamentally limits its ability to deal with network dynamics. Peng Wang 0037, Hong Xu 0001, Zhixiong Niu, Dongsu Han, Yongqiang Xiong |
SoCC | 5 |
| 2016 | ClickNP: Highly flexible and High-performance Network Processing with Reconfigurable HardwareabstractHighly flexible software network functions (NFs) are crucial components to enable multi-tenancy in the clouds. However, software packet processing on a commodity server has limited capacity and induces high latency. While software NFs could scale out using more servers, doing so adds significant cost. This paper focuses on accelerating NFs with programmable hardware, i.e., FPGA, which is now a mature technology and inexpensive for datacenters. However, FPGA is predominately programmed using low-level hardware description languages (HDLs), which are hard to code and difficult to debug. More importantly, HDLs are almost inaccessible for most software programmers. This paper presents ClickNP, a FPGA-accelerated platform for highly flexible and high-performance NFs with commodity servers. ClickNP is highly flexible as it is completely programmable using high-level C-like languages, and exposes a modular programming abstraction that resembles Click Modular Router. ClickNP is also high performance. Our prototype NFs show that they can process traffic at up to 200 million packets per second with ultra-low latency ($< 2\mu$s). Compared to existing software counterparts, with FPGA, ClickNP improves throughput by 10x, while reducing latency by 10x. To the best of our knowledge, ClickNP is the first FPGA-accelerated platform for NFs, written completely in high-level language and achieving 40 Gbps line rate at any packet size. Bojie Li, Kun Tan 0002, Layong Luo, Yanqing Peng, Renqian Luo, Ningyi Xu, Yongqiang Xiong, Peng Cheng 0005 |
SIGCOMM | 7 |
| 2016 | Fast and Cautious: Leveraging Multi-path Diversity for Transport Loss Recovery in Data Centers
Guo Chen 0001, Yuanwei Lu, Yuan Meng 0002, Bojie Li, Kun Tan 0002, Dan Pei, Peng Cheng 0005, Layong Luo, Yongqiang Xiong, Xiaoliang Wang 0001, Youjian Zhao |
USENIX ATC | 9 |
| 2015 | CubicRing: Enabling One-Hop Failure Detection and Recovery for Distributed In-Memory Storage Systems
Yiming Zhang 0003, Chuanxiong Guo, Dongsheng Li 0001, Rui Chu, Yongqiang Xiong |
NSDI | 6 |
| 2014 | Peer-to-peer as an infrastructure service
Jiangchuan Liu, Ke Xu 0002, Yongqiang Xiong, Dongchao Ma, Kai Shuang |
Peer-to-Peer Netw. Appl. | 3 |
| 2013 | Per-packet load-balanced, low-latency routing for clos-based data center networksabstractClos-based networks including Fat-tree and VL2 are being built in data centers, but existing per-flow based routing causes low network utilization and long latency tail. In this paper, by studying the structural properties of Fat-tree and VL2, we propose a per-packet round-robin based routing algorithm called Digit-Reversal Bouncing (DRB). DRB achieves perfect packet interleaving. Our analysis and simulations show that, compared with random-based load-balancing algorithms, DRB results in smaller and bounded queues even when traffic load approaches 100%, and it uses smaller re-sequencing buffer for absorbing out-of-order packet arrivals. Our implementation demonstrates that our design can be readily implemented with commodity switches. Experiments on our testbed, a Fat-tree with 54 servers, confirm our analysis and simulations, and further show that our design handles network failures in 1-2 seconds and has the desirable graceful performance degradation property. Jiaxin Cao, Pengkun Yang, Chuanxiong Guo, Guohan Lu, Yixin Zheng, Yongqiang Xiong, David A. Maltz |
CoNEXT | 9 |
| 2013 | Datacast: A Scalable and Efficient Reliable Group Data Delivery Service for Data CentersabstractReliable Group Data Delivery (RGDD) is a pervasive traffic pattern in data centers. In an RGDD group, a sender needs to reliably deliver a copy of data to all the receivers. Existing solutions either do not scale due to the large number of RGDD groups (e.g., IP multicast) or cannot efficiently use network bandwidth (e.g., end-host overlays). Motivated by recent advances on data center network topology designs (multiple edge-disjoint Steiner trees for RGDD) and innovations on network devices (practical in-network packet caching), we propose Datacast for RGDD. Datacast explores two design spaces: 1) Datacast uses multiple edge-disjoint Steiner trees for data delivery acceleration. 2) Datacast leverages in-network packet caching and introduces a simple soft-state based congestion control algorithm to address the scalability and efficiency issues of RGDD. Our analysis reveals that Datacast congestion control works well with small cache sizes (e.g., 125KB) and causes few duplicate data transmissions (e.g., 1.19%). Both simulations and experiments confirm our theoretical analysis. We also use experiments to compare the performance of Datacast and BitTorrent. In a BCube(4, 1) with 1Gbps links, we use both Datacast and BitTorrent to transmit 4GB data. The link stress of Datacast is 1.01, while it is 1.39 for BitTorrent. By using two Steiner trees, Datacast finishes the transmission in 16.9s, while BitTorrent uses 52s. Jiaxin Cao, Chuanxiong Guo, Guohan Lu, Yongqiang Xiong, Yixin Zheng, Yongguang Zhang, Yibo Zhu 0001, Chen Chen 0019, Ye Tian 0004 |
IEEE J. Sel. Areas Commun. | 4 |
| 2012 | Datacast: a scalable and efficient reliable group data delivery service for data centersabstractReliable Group Data Delivery (RGDD) is a pervasive traffic pattern in data centers. In an RGDD group, a sender needs to reliably deliver a copy of data to all the receivers. Existing solutions either do not scale due to the large number of RGDD groups (e.g., IP multicast) or cannot efficiently use network bandwidth (e.g., end-host overlays). Jiaxin Cao, Chuanxiong Guo, Guohan Lu, Yongqiang Xiong, Yixin Zheng, Yongguang Zhang, Yibo Zhu 0001, Chen Chen 0019 |
CoNEXT | 4 |
| 2012 | Tuning ECN for data center networksabstractThere have been some serious concerns about the TCP performance in data center networks, including the long completion time of short TCP flows in competition with long TCP flows, and the congestion due to TCP incast. In this paper, we show that a properly tuned instant queue length based Explicit Congestion Notification (ECN) at the intermediate switches can alleviate both problems. Compared with previous work, our approach is appealing as it can be supported on current commodity switches with a simple parameter setting and it does not need any modification on ECN protocol at the end servers. Furthermore, we have observed a dilemma in which a higher ECN threshold leads to higher throughput for long flows whereas a lower threshold leads to more senders on incast under buffer pressure. We address this problem with a switch modification only scheme - dequeue marking, for further tuning the instant queue length based ECN to achieve optimal incast performance and long flow throughput with a single threshold value. Our experimental study demonstrates that dequeue marking is effective for increasing the maximum incast senders close to the performance limit of ECN, achieving a gain anywhere from 16% to 140%. Jiabo Ju, Guohan Lu, Chuanxiong Guo, Yongqiang Xiong, Yongguang Zhang |
CoNEXT | 5 |
| 2011 | ServerSwitch: A Programmable and High Performance Platform for Data Center Networks
Guohan Lu, Chuanxiong Guo, Tong Yuan, Yongqiang Xiong, Yongguang Zhang |
NSDI | 7 |
| 2010 | WIND: A scalable and lightweight network topology service for peer-to-peer applicationsabstractWe present an Internet-scale network topology information (NTI) service named WIND for localizing P2P traffic. Central to WIND are the two simple ideas: 1) obtaining NTI directly from routing infrastructures, and 2) leveraging existing, widely deployed DNS caches for NTI delivery. WIND fulfills the fidelity, flexibility and scalability requirement of an effective NTI service. WIND is deployed in CERNET. We conduct extensive trace-driven emulations on PlanetLab. Experimental results confirm the effectiveness of the WIND service. Hongqiang Liu, Yongqiang Xiong, CongXiao Bao, Xing Li 0001, Guobin Shen, Dan Li 0001 |
NOMS | 2 |
| 2010 | mTreebone: A Collaborative Tree-Mesh Overlay Network for Multicast Video StreamingabstractRecently, application-layer overlay networks have been suggested as a promising solution for live video streaming over the Internet. To organize a multicast overlay, a natural structure is a tree, which, however, is known vulnerable to end-hosts dynamics. Data-driven approaches address this problem by employing a mesh structure, which enables data exchanges among multiple neighbors, and thus, greatly improves the overlay resilience. It unfortunately suffers from an efficiency-delay trade-off, because data have to be pulled from mesh neighbors by using extra notifications periodically. In this paper, we closely examine the contributions of overlay nodes, and argue that performance of a mesh overlay closely depends on a small set of stable backbone nodes. This is validated through a real trace study on PPLive, the largest commercial application-layer live streaming system to date. Motivated by this observation, we then suggest a novel collaborative tree-mesh design that leverages both mesh and tree structures. The key idea is to identify a set of stable nodes to construct a tree-based backbone, called treebone, with most of the data being pushed over this backbone. These stable nodes, together with others, are further organized through an auxiliary mesh overlay, which facilitates the treebone to accommodate node dynamics and fully exploit the available bandwidth between overlay nodes. This hybrid design, referred to as mTreebone, brings a series of unique and critical design challenges. In particular, the identification of stable nodes and seamless data delivery using both push and pull methods. In this paper, we present optimized solutions to these problems, which reconcile the two overlays under a coherent framework with controlled overhead. We evaluate mTreebone through both simulations and PlanetLab experiments. The results demonstrate the superior efficiency and robustness of this hybrid solution in both static and dynamic scenarios. Feng Wang 0001, Yongqiang Xiong, Jiangchuan Liu |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2009 | Experimental Study of Broadcatching in BitTorrentabstractBroadcatching is a promising mechanism to improve the experience of BitTorrent users by automatically downloading files advertised through RSS feeds. However, though widely used, the mechanism itself has not been well studied. In this paper, we conducted extensive experiments on PlanetLab to evaluate the performance of Broadcatching under different typical scenarios. The results demonstrated the effectiveness of the broadcatching: it reduces the average completion time and downloading failure ratio. It also improves the overall fairness of the system: the subscribers are encouraged to share more while downloading faster, which results in the increased share ratio. Our study is the first work to systematically evaluate the benefit of broadcatching and sheds lights on how to improve performance of BitTorrrent by manipulating peer's behavior like Broadcatching. Zengbin Zhang, Yang Chen 0001, Yongqiang Xiong, Guobin Shen, Hongqiang Liu, Beixing Deng, Xing Li 0001 |
CCNC | 4 |
| 2009 | Pharos: accurate and decentralised network coordinate systemabstractNetwork coordinates (NC) system is an efficient mechanism for Internet distance prediction with scalable measurements. The intrinsical cause for the unsatisfactory accuracy of the simulation-based NC algorithms has been identified. Then Pharos, a fully decentralised and hierarchical scheme, is proposed to solve this problem. Pharos leverages multiple coordinate sets at different distance scales, with the right scale being chosen for prediction each time. We evaluate the performance of Pharos system with the King data set and latency data from PlanetLab, and compare it with the representative NC system, Vivaldi. The experimental results show that Pharos greatly outperforms Vivaldi in Internet distance prediction without adding any significant overhead. Our extensive evaluation results also demonstrate that Pharos can significantly improve the performance in distributed Internet applications, such as overlay multicast and server selection. Yang Chen 0001, Yongqiang Xiong, Xiaohui Shi, Jiwen Zhu, Beixing Deng, Xing Li 0001 |
IET Commun. | 2 |
| 2009 | Optimal Prefetching Scheme in P2P VoD Applications With Guided SeeksabstractMost existing peer-to-peer (P2P) video-on-demand (VoD) systems have been designed and optimized for the sequential playback. In practice, users often want to seek to the positions they are interested in. Such frequent seeks raise greater challenges to the design of the prefetching scheme. In this work, we first propose the concept of guided seeks. With the guidance, users can perform more efficient seeks to the desired positions. The guidance can be obtained from collective seeking statistics of other peers who have watched the same title in the previous and/or concurrent sessions. However, it is very challenging to aggregate the statistics efficiently, timely and in a completely distributed way. We design the hybrid sketches that not only capture the seeking statistics at significantly reduced space and time complexity, but also adapt to the popularity of the video. From the collected seeking statistics, we estimate the segment access probability, based on which we further develop an optimal prefetching scheme and an optimal cache replacement policy to minimize the expected seeking delay at every viewing position. Through extensive simulations, we demonstrate that the proposed prefetching framework significantly reduces the seeking delay compared to the sequential prefetching scheme. Guobin Shen, Yongqiang Xiong, Ling Guan |
IEEE Trans. Multim. | 3 |
| 2009 | Optimizing the Throughput of Data-Driven Peer-to-Peer StreamingabstractDuring recent years, the Internet has witnessed a rapid growth in deployment of data-driven (or swarming based) peer-to-peer (P2P) media streaming. In these applications, each node independently selects some other nodes as its neighbors (i.e. gossip-style overlay construction), and exchanges streaming data with the neighbors (i.e. data scheduling). To improve the performance of such protocol, many existing works focus on the gossip-style overlay construction issue. However, few of them concentrate on optimizing the streaming data scheduling to maximize the throughput of a constructed overlay. In this paper, we analytically study the scheduling problem in data-driven streaming system and model it as a classical min-cost network flow problem. We then propose both the global optimal scheduling scheme and distributed heuristic algorithm to optimize the system throughput. Furthermore, we introduce layered video coding into data-driven protocol and extend our algorithm to deal with the end-host heterogeneity. The results of simulation with the real world traces indicate that our distributed algorithm significantly outperforms conventional ad hoc scheduling strategies especially in stringent buffer and bandwidth constraints. Meng Zhang 0001, Yongqiang Xiong, Qian Zhang 0001, Lifeng Sun, Shiqiang Yang |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2008 | Stable Peers: Existence, Importance, and Application in Peer-to-Peer Live Video StreamingabstractThis paper presents a systematic in-depth study on the existence, importance, and application of stable nodes in peer- to-peer live video streaming. Using traces from a real large-scale system as well as analytical models, we show that, while the number of stable nodes is small throughout a whole session, their longer lifespans make them constitute a significant portion in a per-snapshot view of a peer-to-peer overlay. As a result, they have substantially affected the performance of the overall system. Inspired by this, we propose a tiered overlay design, with stable nodes being organized into a tier-1 backbone for serving tier-2 nodes. It offers a highly cost-effective and deployable alternative to proxy-assisted designs. We develop a comprehensive set of algorithms for stable node identification and organization. Specifically, we present a novel structure,LabeledTree, for the tier-1 overlay, which, leveraging stable peers, simultaneously achieves low overhead and high transmission reliability. Our tiered framework flexibly accommodates diverse existing overlay structures in the second tier. Our extensive simulation results demonstrated that the customized optimization using selected stable nodes boosts the streaming quality and also effectively reduces the control overhead. This is further validated through prototype experiments over the PlanetLab network. Feng Wang 0001, Jiangchuan Liu, Yongqiang Xiong |
INFOCOM | 3 |
| 2008 | Probabilistic prefetching scheme for P2P VoD applications with frequent seeksabstractIn Peer-to-Peer Video-on-Demand (P2P VoD) applications, users tend to seek to the positions that they are interested in. The frequent seeks raise a great challenge to the design of the prefetching scheme. In this paper, we propose a probabilistic prefetching framework to reduce the seeking distance. Each peer performs prefetching based on the segment access probability, which is estimated from the seeking statistics in the previous sessions. It is a challenging task to collect the seeking statistics in a distributed P2P network. In the proposed framework, we employ FM sketches to represent the seeking statistics, thus greatly reducing the space and time complexity. The simulation results show that the proposed prefetching scheme can approach closer to the desired seeking positions compared to the prefetching scheme neglecting the user viewing pattern. Guobin Shen, Yongqiang Xiong, Ling Guan |
ISCAS | 3 |
| 2007 | Pharos: A Decentralized and Hierarchical Network Coordinate System for Internet Distance PredictionabstractNetwork coordinates (NC) system is an efficient mechanism for Internet distance prediction with limited measurements. In this paper, we identify the intrinsical cause for the inadequate accuracy of the simulation based NC algorithms. We then propose Pharos, a fully decentralized and hierarchical scheme, to remedy this problem. Pharos leverages multiple coordinate sets at different distance scales, with the right scale being chosen for prediction each time. We evaluate the performance of Pharos system with the King data set and latency data from PlanetLab, and compare it with the representative NC system, Vivaldi. The experimental results show that Pharos outperforms Vivaldi much without adding any significant overhead. Yang Chen 0001, Yongqiang Xiong, Xiaohui Shi, Beixing Deng, Xing Li 0001 |
GLOBECOM | 2 |
| 2007 | mTreebone: A Hybrid Tree/Mesh Overlay for Application-Layer Live Video MulticastabstractApplication-layer overlay networks have recently emerged as a promising solution for live media multicast on the Internet. A tree is probably the most natural structure for a multicast overlay, but is vulnerable in the presence of dynamic end-hosts. Data-driven approaches form a mesh out of overlay nodes to exchange data, which greatly enhances the resilience. It however suffers from an efficiency-latency tradeoff, given that the data have to be pulled from mesh neighbors with periodical notifications. In this paper, we suggest a novel hybrid tree/mesh design that leverages both overlays. The key idea is to identify a set of stable nodes to construct a tree-based backbone, called treebone, with most of the data being pushed over this backbone. These stable nodes, together with others, are further organized through an auxiliary mesh overlay, which facilitates the treebone to accommodate node dynamics and fully exploit the available bandwidth between overlay nodes. This hybrid design, referred to as mTreebone, is braced by our real trace studies, which show strong evidence that the performance of an overlay closely depends on a small set of backbone nodes. It however poses a series of unique and critical design challenges, in particular, the identification of stable nodes and seamless data delivery using both push and pull methods. In this paper, we present optimized solutions to these problems, which reconcile the two overlays under a coherent framework with controlled overhead. We evaluate mTreebone through both simulations and PlanetLab experiments. The results demonstrate the superior efficiency and robustness of this hybrid solution. Feng Wang 0001, Yongqiang Xiong, Jiangchuan Liu |
ICDCS | 2 |
| 2007 | Optimizing the Throughput of Data-Driven Based Streaming in Heterogeneous Overlay Network
Meng Zhang 0001, Chunxiao Chen, Yongqiang Xiong, Qian Zhang 0001, Shiqiang Yang |
MMM (1) | 3 |
| 2007 | Cross-layer optimization in ultra wideband networks
Qi Wu 0002, Jingping Bi, Zihua Guo, Yongqiang Xiong, Qian Zhang 0001, Zhongcheng Li |
Sci. China Ser. F Inf. Sci. | 4 |
| 2006 | Anycast Routing in Delay Tolerant NetworksabstractAnycast routing is very useful for many applications such as resource discovery in delay tolerant networks (DTNs). In this paper, based on a new DTN model, we first analyze the any-cast semantics for DTNs. Then we present a novel metric named EMDDA (expected multi-destination delay for anycast) and a corresponding routing algorithm for anycast routing in DTNs. Extensive simulation results show that the proposed EMDDA routing scheme can effectively improve the efficiency of anycast routing in DTNs. It outperforms another algorithm, minimum expected delay (MED) algorithm, by 11.3% on average in term of routing delays and by 19.2% in term of average max queue length. Yili Gong, Yongqiang Xiong, Qian Zhang 0001, Zhensheng Zhang, Wenjie Wang 0006, Zhiwei Xu 0002 |
GLOBECOM | 2 |
| 2006 | On the Optimal Scheduling for Media Streaming in Data-driven Overlay NetworksabstractThe Internet has witnessed a rapid growth in deployment of data-driven overlay network (DON) based streaming applications during recent years. In these applications, each node independently selects some other nodes as its neighbors (i.e. overlay construction), and exchanges streaming data with these neighbors (i.e. data scheduling). This scheme improves the robustness of the system. However, most of the work in the literature focused on the construction problem, and very few addressed its scheduling problem which is also very important for the overall performance. In this paper, we analytically study the scheduling problem in DON and model it as a classical min-cost network flow problem. We then propose both the global optimal scheduling scheme and distributed heuristic algorithm to maximize the system throughput. Experimental results indicate that our algorithms outperform other schemes and the throughput gain is up to 80%. Meng Zhang 0001, Yongqiang Xiong, Qian Zhang 0001, Shiqiang Yang |
GLOBECOM | 2 |
| 2006 | Detecting Malicious Hosts in the Presence of Lying Hosts in Peer-to-Peer StreamingabstractCurrent peer-to-peer (P2P) streaming systems often assume that hosts are cooperative. However, this may not be true in the open environment of the internet. In this paper, we discuss how to detect malicious hosts (e.g., with attacking actions and abnormal behavior) based on their history performance. In our system, each host monitors the performance of its neighbor(s) and reports this to a server. Based on the reports, the server computes host reputation with hosts of low reputation being malicious. A problem is that hosts may lie by submitting forged reports to the server. We hence formulate the reputation computing problem in the presence of lying hosts as a minimization problem and solve it by the traditional Levenberg-Marquardt algorithm. Simulation results show that our scheme can efficiently detect malicious hosts with high accuracy Shueng-Han Gary Chan, Wai-Pun Ken Yiu, Yongqiang Xiong, Qian Zhang 0001 |
ICME | 4 |
| 2006 | Ripple-Stream: Safeguarding P2P Streaming Against Dos AttacksabstractCompared with file-sharing and distributed hash table (DHT) network, P2P video streaming is more vulnerable to denial of service (DoS) attacks because of its high bandwidth demand and stringent time requirement. This paper studies the design of DoS resilient streaming networks using credit systems. We propose a novel framework-ripple-stream-to improve DoS resilience of P2P streaming. Ripple-stream leverages existing credit systems to introduce credit constraints in overlay construction such that malicious nodes are pushed to the fringe of overlays. Combining credit constraints with overlay optimization techniques, ripple-stream can achieve both DoS resilience and overlay efficiency Wenjie Wang 0006, Yongqiang Xiong, Qian Zhang 0001, Sugih Jamin |
ICME | 2 |
| 2006 | Prediction-based routing for real time communications in wireless multi-hop networksabstractReal time communication (RTC) has critical quality of service (QoS) requirements, which is much more challenging in wireless multi-hop networks. Traditional measurement-based routing schemes often ignore the interference from the coming RTC traffic itself (i.e. self-traffic), so they can not get an accurate quality estimation of the path to serve the coming RTC traffic. In this paper, we propose a novel prediction-based routing metric, PPTT (Path Predicted Transmission Time), to estimate end-to-end delay of RTC traffics. PPTT is traffic-aware by taking explicit consideration of both self-traffic and neighboring traffics interfering with the RTC flow, and thus offers an accurate prediction of transmission delay. By selecting route with minimal PPTT, the quality of service for the coming RTC flow will be improved, in terms of end-to-end delay and goodput. To evaluate the performance, we implement PPTT scheme and study its performance in a wireless multi-hop test bed consisting of 32 nodes equipped with IEEE 802.11 a/b/g combo cards, and we also conduct extensive simulations with different random topologies in network simulator NS2 for a more comprehensive comparison. Experiment results show that this routing metric outperforms other non prediction-based routing metric such as ETX (Expected Transmission Count) and WCETT (Weighted Cumulative Expected Transmission Time) in terms of delay and goodput in wireless multi-hop networks. Shouyi Yin, Yongqiang Xiong, Qian Zhang 0001, Xiaokang Lin |
QSHINE | 2 |
| 2006 | Joint routing and topology formation in multihop UWB networksabstractThis paper addresses the throughput optimization problem in multihop ultra-wideband (UWB) networks by jointly considering network topology formation and routing. Given a spatial distribution of UWB devices and traffic requirement, we want to form piconets and select paths to maximize the network throughput. Although there have been several works considering the problem of selecting paths to achieve the optimal throughput in multihop wireless networks, to the best of our knowledge, none of them takes the topology formation into the consideration. In this paper, we use Boolean matrices to model role assignment in UWB networks and formulate the throughput optimization problem as a nonlinear programming (NLP) problem. Since the throughput optimization problem is NP-hard, we give an upper bound of the optimal throughput by relaxing some constraints and using pseudo-Boolean optimization to linearize the NLP. We prove that the solution of the upper bound is at most three times of the optimal throughput. Based on the topology formed by solving the upper bound, we formulate a lower bound of the optimal throughput as a linear programming problem and use column generation to solve the lower bound. Numerical results show that the lower bound is very close to the upper bound. Simulation results demonstrate the effectiveness of the scheme. Qi Wu 0002, Yongqiang Xiong, Qian Zhang 0001, Zihua Guo, Xiang-Gen Xia 0001, Zhongcheng Li |
IEEE J. Sel. Areas Commun. | 2 |
| 2006 | Traffic-aware routing for real-time communications in wireless multi-hop networksabstractAbstract In this paper, we propose a novel traffic‐aware routing metric for real‐time communications (RTC) in wireless multi‐hop networks. Our routing metric, path predicted transmission time (PPTT), is designed to choose a high‐quality path for RTC flow between a source and a destination. PPTT can serve as both single‐radio and multi‐radio routing metric for RTC flow. RTC has critical quality of service (QoS) requirements in terms of delay, bandwidth and so on. Traditional measurement‐based routing schemes ignore the interference from the coming RTC flow itself (i.e. self‐traffic), so they may choose the inefficient path to serve the coming RTC flow due to the inaccurate quality estimation of the transmission path. PPTT takes explicit consideration of both self‐traffic and neighbouring traffic interfering with the RTC flow, and thus offers an accurate estimation of path transmission delay. Through differentiating the links by the wireless channel/radio they are using, PPTT has the capability to choose a high‐quality path for the coming RTC flow in both single‐radio and multi‐radio networks. To evaluate the performance, we implement PPTT scheme and study its performance in a wireless multi‐hop testbed consisting of 32 nodes equipped with two IEEE 802.11a/b/g combo cards, and we also conduct extensive simulations with different random topologies in network simulator NS‐2 for a more comprehensive comparison. The results of simulation and experiment show that this routing metric outperforms other non‐traffic‐aware one such as expected transmission count (ETX) and weighted cumulative expected transmission time (WCETT) in terms of delay and goodput in both single‐radio and multi‐radio wireless networks. Copyright © 2006 John Wiley & Sons, Ltd. Shouyi Yin, Yongqiang Xiong, Qian Zhang 0001, Xiaokang Lin |
Wirel. Commun. Mob. Comput. | 2 |
| 2005 | An experimental study on multi-channel multi-radio multi-hop wireless networksabstractIn this paper, we present experimental studies on Multi-channel Multi-radio Multi-hop (M3) wireless networks. The interference of multiple channels and multiple radios is studied. The result may help to understand effect of multi-channel and multi-radio in real system and is useful as reference for network performance optimization. We evaluate the coexistence of real-time traffic and best-effort traffic in multi-hop wireless networks. Our experiments show that coordination among wireless nodes is critical for optimal network performance, especially in the scenarios with quality of service (QoS) requirements. To perform the experiments, an M3 testbed with 32 nodes is built. The testbed is based on Windows operating system and provides convenient programming interface and management tools. The architecture and key modules of the testbed design are described. Yongqiang Xiong, Yang Yang 0022, Pengzhi Xu, Qian Zhang 0001 |
GLOBECOM | 2 |
| 2005 | Interference aware metric for dense multi-hop wireless networksabstractA key issue impacting the performance of multi-hop wireless networks is wireless interference among neighboring nodes. In this paper, we study the impact of interference on mean delay and available bandwidth residing at wireless nodes and present a novel interference aware metric, named network allocation vector count (NAVC). The design of NAVC as a metric for the AODV routing protocols, as well as a metric for transmit power control are described in detail. Our simulations demonstrate the poor performance of minimum hop-counts routing protocols, and confirm that NAVC based routing protocol can greatly improve performance. The average throughput increases by up to 29% for UDP CBR traffic. For scenarios of densely deployed nodes, the throughput improvement is often a factor near two, suggesting that NAVC will become more useful as networks grow larger and paths become longer. These approaches are essential for emerging applications such as sensor networks where interference is heavy and bandwidth is limited. Liran Ma, Qian Zhang 0001, Yongqiang Xiong, Wenwu Zhu 0001 |
ICC | 3 |
| 2005 | Robust and efficient path diversity in application-layer multicast for video streamingabstractApplication-layer multicast (ALM), as alternative to IP multicast, provides group communication without the need for network infrastructure support. To improve the reliability of ALM service, path diversity has been studied and two schemes to construct diverse paths for hosts are proposed. One is the random multicast forest (RMF) and the other is topology-aware hierarchical arrangement graph (THAG). RMF makes the paths from the media source to a participating host diverse by selecting parents for each host randomly, while THAG makes the paths node-disjoint by constructing multiple independent multicast trees, where any interior node in a multicast tree will be leaf node in all the other multicast trees. Topology-awareness is implemented in both schemes to make them efficient for media delivery. We compare the reliability and efficiency of THAG and RMF through extensive simulation. The results show that the reliability of THAG has been improved up to 20% compared with RMF. The efficiency metrics, such as relative delay penalty, link stress, and delay variation among different trees in THAG, are also smaller than or almost the same as that in RMF. The results indicate that THAG is a reliable and efficient ALM scheme for streaming media service. Ruixiong Tian, Qian Zhang 0001, Zhe Xiang, Yongqiang Xiong, Xing Li 0001, Wenwu Zhu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |