Xiaoliang Wang 0001

dblp:02/3450-1 · DBLP profile ↗
← Back
98ranked-venue papers
12as first author
48since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 57 · 5 first-author · 31 since 2021Systems, architecture and hardware · 21 · 12 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 Rethinking MoE Routing for Commodity GPUs
abstract
Commodity GPUs make large-model serving economically attractive, but they are typically connected via PCIe rather than high-bandwidth GPU fabrics. This is a poor fit for expert-parallel Mixture-of-Experts (MoE) inference, where each MoE layer routes tokens to selected experts via all-to-all communication, and PCIe crossings can become a major latency bottleneck. Existing serving systems and communication libraries optimize this traffic but treat router-selected expert assignments as fixed. We propose PCIe-MoE, an inference-time mechanism that decides when a low-confidence remote expert is worth a costly PCIe hop, and when a nearby, score-similar expert can be substituted instead. PCIe-MoE combines topology profiling, replaceability-aware placement, risk-aware remapping, and quality guards to reshape token-to-expert routing without retraining.
Meng Li 0010, Siyuan Tong, Qingkai Meng 0001, Xiaoliang Wang 0001, Haipeng Dai 0001
APNet5
2026 SemDNS: A Declarative Semantics for DNS Resolution
Kaiqiang Hu, Haizhou Du, Xiaoliang Wang 0001
SIGCOMM4
2026 PSN-PATH: When Multipath RDMA Meets Lossy Networks
Zhexiong Li, Shugui Wei, Puyu Zhao, Yuepeng Li, Lin Gu 0002, Deze Zeng, Xiaoliang Wang 0001, Laiping Zhao
SIGCOMM8
2026 GPU-Centric Stateless LLM Serving With GIGANETS
abstract
Giganetes is a GPU-centric architecture for stateless LLM serving that externalizes KV caches to a disaggregated remote memory pool via GPUDirect RDMA. By treating remote memory as a GPU-addressable tier via GPUDirect RDMA, Giganetes eliminates session affinity constraints: any GPU can serve any request, enabling near-linear horizontal scaling in a Kubernetes-native deployment. A Scatter/Gather I/O interface bypasses the CPU and host memory entirely, achieving 52.4 GB/s application-level read throughput on our 4×200 Gbps RDMA testbed. A session-level metadata abstraction and proactive readahead mechanism reduce GPU bubbles by overlapping remote KV fetches with prefill computation and scheduling slack. On a 4-node H800 cluster, Giganetes delivers 33% higher throughput (QPS 2.4 vs. 1.8) and 1.75× lower P95 TPOT than PD-disaggregation with sticky sessions, with the gain driven by scheduling flexibility rather than faster transport alone.
Xiaoliang Wang 0001, Zhenwei Pi, Cam-Tu Nguyen
SIGCOMM1
2026 DTCC: Decision Transformer-driven framework for adaptive network congestion control
abstract
Existing learning-based congestion control methods suffer from myopic decision-making due to their reliance on single-timestep states and fail to model long-term dependencies due to architectural constraints (e.g., recurrent networks’ vanishing gradients). To address these issues, we propose a Decision Transformer-based network congestion control framework named DTCC. DTCC is the first to unify long-context modeling and real-time decision-making within a 4-layer autoregressive Transformer, replacing traditional Markov decision paradigms with sequence-to-action mapping. With enhancement learning strategy such as stochasticity-aware training, DTCC achieves efficient and generalizable performance from heterogeneous dataset. Extensive experiments demonstrate DTCC’s supremacy: it achieves 16.67–29.55% higher winning rate compared to state-of-the-art baselines (e.g., Sage) across diverse network scenarios and 8.33%–29.17% higher winning rate under unseen highly variable network. Leveraging a lightweight Transformer, DTCC enables real-time deployment with approximately 2.8 ms inference per step on general CPU devices. To the best of our knowledge, this is the first work to employ Decision Transformer for training an intelligent congestion control mechanism. Our work, therefore, showcases the potential of combining reinforcement learning with advanced Transformer architectures in real-time network control.
Xiaolan Ji, Biao Han 0003, Xiaoliang Wang 0001, Ruidong Li 0001, Jinshu Su
Comput. Networks3
2026 Fine-Grained Scheduling of In-Network Aggregation Resources for Efficient Machine Learning Service
Shichen Dong, Zhixiong Niu, Mingchao Zhang, Zhiying Xu, Chuntao Hu, Pengzhi Zhu, Qingchun Song, Peng Cheng 0005, Cam-Tu Nguyen, Shaoling Sun, Xiaohu Xu, Yongqiang Xiong, Wei Wang 0002, Xiaoliang Wang 0001, Guihai Chen
IEEE Trans. Netw.15
2026 Analysis of Pyrrha: Congestion-Root-Based Flow Control Is Most Cost-Effective to Eliminate Head-of-Line Blocking
abstract
In modern datacenters, the effectiveness of end-to-end congestion control (CC) is quickly diminishing with the rapid bandwidth evolution. Per-hop flow control (FC) can react to congestion more promptly. However, a coarse-grained FC can result in Head-Of-Line (HOL) blocking. A fine-grained, per-flow FC can eliminate HOL blocking caused by flow control, however, it does not scale well. This paper presents Pyrrha, a scalable flow control approach that provably eliminates HOL blocking while using a minimum number of queues. In Pyrrha, flow control first takes effect on the root of the congestion, i.e., the port where congestion occurs. And then flows are controlled according to their contributed congestion roots. A prototype of Pyrrha is implemented on Tofino2 switches. Compared with state-of-the-art approaches, the average FCT of uncongested flows is reduced by 42%-98%, and 99th-tail latency can be$1.6\times $-$215\times $lower, without compromising the performance of congested flows.
Zhaochen Zhang, Peirui Cao, Chang Liu 0001, Yizhi Wang 0004, Vamsi Addanki, Stefan Schmid 0001, Qingyue Wang, Xiaoliang Wang 0001, Jiaqi Zheng 0001, Tao Wu 0011, Bingyang Liu, Wan-Chun Dou, Guihai Chen, Chen Tian 0001, Fu Xiao 0001
IEEE Trans. Netw.10
2025 Using Analytical Performance/Power Model and Fine-Grained DVFS to Enhance AI Accelerator Energy Efficiency
abstract
Recent advancements in deep learning have significantly increased AI processors' energy consumption, which is becoming a critical factor limiting AI development. Dynamic Voltage and Frequency Scaling (DVFS) stands as a key method in power optimization. However, due to the latency of DVFS control in AI processors, previous works typically apply DVFS control at the granularity of a program's entire duration or sub-phases, rather than at the level of AI operators.
Yijia Zhang 0002, Fuchun Wei, Bingqiang Wang, Yanlin Liu, Zhiheng Hu, Xiaoxin Xu, Xiaoliang Wang 0001, Wan-Chun Dou, Guihai Chen, Chen Tian 0001
ASPLOS (1)10
2025 Marlin: Enabling High-Throughput Congestion Control Testing in Large-Scale Networks
abstract
Cloud providers require high-throughput traffic to test the effectiveness of congestion control (CC) configurations (i.e., CC algorithm selection and their parameter settings) in networks. A network tester capable of evaluating CC configurations needs to fulfill the following requirements: (R1) Capable of generating traffic with CC behaviors. (R2) Ability to customize CC algorithms. (R3) High throughput CC traffic generation. However, existing network testers fail to meet these requirements simultaneously. The paper presents Marlin, a novel high-throughput network tester designed for CC evaluation. Marlin leverages a high-throughput, low-programmability device to amplify the traffic generated by a low-throughput, high-programmability device. The low-throughput device is responsible for complex computational tasks, such as running CC and flow scheduling algorithms, and communicates with the high-throughput device at a high frequency using small packets to instruct it to generate high-throughput traffic with CC behaviors. This hybrid approach allows for customizable, high-throughput CC testing. Our experiments demonstrate that Marlin can accurately emulate CC behaviors and replicate real-world scenarios. Marlin can generate 1.2 Tbps of CC traffic using a single programmable switch pipeline and one 100 Gbps port of an FPGA NIC, supporting up to 65,536 concurrent flows.
Li Wang 0110, Jingzhi Wang, Songyue Liu, Keqiang He, Jian Wang 0025, Xiaoliang Wang 0001, Wan-Chun Dou, Guihai Chen, Chen Tian 0001
EuroSys7
2025 Enabling Virtual Priority in Data Center Congestion Control
abstract
In data center networks, various types of traffic with strict performance requirements operate simultaneously, necessitating effective isolation and scheduling through priority queues. However, most switches support only around ten priority queues. Virtual priority can address this limitation by emulating multi-priority queues on a single physical queue, but existing solutions often require complex switch-level scheduling and hardware changes. Our key insight is that virtual priority can be achieved by carefully managing bandwidth contention in a physical queue, which is traditionally handled by congestion control (CC) algorithms. Hence, the virtual priority mechanism needs to be tightly coupled with CC. In this paper, we propose PrioPlus, a CC enhancement algorithm that can be integrated with existing congestion control schemes to enable virtual priority transmission. PrioPlus assigns specific delay ranges to different priority levels, ensuring that flows transmit only when the delay is within the assigned range, effectively meeting virtual priority requirements. Compared to Swift CC with physical priority queues, PrioPlus provides strict priority for high-priority flows without impacting performance sensibly. Meanwhile, it benefits low-priority flows from 25% to 41% as its priority-aware design enhances CC's ability to fully utilize available bandwidth once higher-priority traffic completes. As a result, in coflow and model training scenarios, PrioPlus improves job completion times by 21% and 33%, respectively, compared to Swift with physical priority queues.
Zhaochen Zhang, Feiyang Xue, Keqiang He, Zhimeng Yin 0001, Gianni Antichi, Yizhi Wang 0004, Rui Ning, Haixin Nan, Xu Zhang 0006, Peirui Cao, Xiaoliang Wang 0001, Wan-Chun Dou, Guihai Chen, Chen Tian 0001
EuroSys12
2025 Mina: Fine-Grained In-network Aggregation Resource Scheduling for Machine Learning Service
Shichen Dong, Zhixiong Niu, Mingchao Zhang, Zhiying Xu, Chuntao Hu, Pengzhi Zhu, Qingchun Song, Peng Cheng 0005, Cam-Tu Nguyen, Shaoling Sun, Xiaohu Xu, Yongqiang Xiong, Wei Wang 0002, Xiaoliang Wang 0001
INFOCOM15
2025 SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
abstract
KV cache eviction has emerged as an effective solution to alleviate resource constraints faced by LLMs in long-context scenarios. However, existing token-level eviction methods often overlook two critical aspects: (1) their irreversible eviction strategy fails to adapt to dynamic attention patterns during decoding (the saliency shift problem), and (2) they treat both marginally important tokens and truly unimportant tokens uniformly, despite the collective significance of marginal tokens to model performance (the marginal information over-compression problem). To address these issues, we design two compensation mechanisms based on the high similarity of attention matrices between LLMs with different scales. We propose SmallKV, a small model assisted compensation method for KV cache compression. SmallKV can maintain attention matching between different-scale LLMs to: 1) assist the larger model in perceiving globally important information of attention; and 2) use the smaller model’s attention scores to approximate those of marginal tokens in the larger model. Extensive experiments on benchmarks including GSM8K, BBH, MT-Bench, and LongBench demonstrate the effectiveness of SmallKV. Moreover, efficiency evaluations show that SmallKV achieves 1.75 - 2.56 times higher throughput than baseline methods, highlighting its potential for efficient and performant LLM inference in resource constrained environments.
Yajuan Peng, Cam-Tu Nguyen, Zuchao Li, Xiaoliang Wang 0001, Hai Zhao 0001, Xiaoming Fu 0001
NeurIPS5
2025 Pyrrha: Congestion-Root-Based Flow Control to Eliminate Head-of-Line Blocking in Datacenter
Zhaochen Zhang, Chang Liu 0001, Yizhi Wang 0004, Vamsi Addanki, Stefan Schmid 0001, Qingyue Wang, Xiaoliang Wang 0001, Jiaqi Zheng 0001, Tao Wu 0011, Bingyang Liu, Wan-Chun Dou, Guihai Chen, Chen Tian 0001
NSDI9
2025 When P4 Meets Run-to-completion Architecture
Jiaqi Zheng 0001, Xiaoliang Wang 0001, Luyou He, Xiaofei Lai, Fuguang Huang, Wan-Chun Dou, Guihai Chen, Chen Tian 0001
NSDI5
2025 Barre: Empowering Simplified and Versatile Programmable Congestion Control in High-Speed AI Clusters
Yajuan Peng, Xiaolong Zhong, Haohan Xu, Zhuo Jiang, Jianxi Ye, Xiaoliang Wang 0001, Xiaoming Fu 0001, Huichen Dai
USENIX ATC10
2025 Accelerating Model Training on Ascend Chips: An Industrial System for Profiling, Analysis and Optimization
Zhibin Wang 0002, Ruyi Zhang 0005, Chen Tian 0001, Xiaoliang Wang 0001, Wan-Chun Dou, Guihai Chen, Bingqiang Wang, Yonghong Tian 0001, Yan Zhang 0002, Hui Wang 0030, Fuchun Wei, Boquan Sun, Bin She, Teng Su, Yaoyuan Wang, Guyue Liu
USENIX ATC6
2025 Reliability-aware hybrid SFC backup and deployment in edge computing
Yue Zeng 0002, Shanshan Lin, Bin Tang 0002, Xiaoliang Wang 0001, Zhihao Qu, Song Guo 0001, Junlong Zhou
Comput. Networks5
2025 Troubleshooting Programmable Data Planes via Real-Time Table Information Recording
abstract
While the flexibility of programmable switches brings opportunities, it also introduces security risks. Hence, it is vital to conduct effective troubleshooting in the programmable switch to mitigate frequent network failures. However, troubleshooting programmable switch failures is challenging due to their enhanced flexibility and functionality compared to regular switches, posing increased difficulty in debugging, particularly with limited debugging tools and information. To address this problem, we propose an efficient troubleshooting method that records real-time information about packets in the data plane, including the tables involved in packet processing. Unfortunately, due to hardware limitations, it is infeasible to record all tables’ information in the data plane. Thus, the key is to find the table set reflecting the execution path a packet goes through while minimizing the resource overhead. We first represent P4 programs as a probabilistic transition directed acyclic graph (DAG) and employ information entropy to quantify the information within a set of tracked tables. Then, we adopt a two-step approach and design algorithms to find both optimal and approximately optimal table record plans. The evaluation results show the efficacy of the proposed method, including achieving the same path recovery rate as the related works with less than one-third of the resource consumption.
Chengyuan Huang, Yibo Xiao, Tianfan Zhang, Bingheng Yan, Ahmed M. Abdelmoniem, Gianni Antichi, Xiaoliang Wang 0001, Fu Xiao 0001, Wan-Chun Dou, Guihai Chen, Chen Tian 0001
IEEE Trans. Netw.9
2025 An Anatomy of Token-Based Congestion Control
abstract
Congestion control protocols play a vital role in enhancing the performance of various applications within datacenter networks. While reactive congestion control (RCC) protocols are widely deployed in commercial datacenters, the research community has actively explored token-based proactive congestion control (TCC) protocols to further push the boundaries of performance. However, despite the emergence of numerous TCC variants, there has been a lack of systematic exploration in the design space of TCC. This paper aims to bridge this gap by proposing a framework for understanding the design choices within the TCC approach. In this study, we systematically analyze different design choices of TCC approaches and leverage this understanding to develop a novel TCC protocol called ToCC. To implement ToCC, we address a set of challenges and deploy it in NP-based smart NICs. We compare ToCC with state-of-the-art TCC and RCC protocols through extensive large-scale simulations and testbed evaluations. The results demonstrate that ToCC exhibits robustness in achieving low latency across various scenarios. Additionally, ToCC effectively reduces buffer occupancy by 4.8 times compared to existing approaches, and under incast scenarios, it significantly shortens flow completion time by up to 90%. Congestion control protocols are crucial for optimizing the performance of datacenter network applications. Although reactive congestion control (RCC) protocols are commonly used in commercial datacenters, researchers have been exploring token-based proactive congestion control (TCC) protocols to further enhance network performance. Despite the development of numerous TCC variants, there has not been a thorough examination of the design space of TCC protocols until now. This paper aims to address this gap by introducing a framework for understanding the design choices within the TCC approach for TCC protocols. By analyzing various design aspects of TCC approaches, we create a novel TCC protocol called ToCC. At the central of ToCC design is that it leverages congestion control mechanisms over tokens. To implement ToCC, we tackle several challenges and integrate it into NP-based smart NICs. Comparing ToCC with state-of-the-art TCC and RCC protocols through extensive large-scale simulations and testbed evaluations, we find that ToCC consistently achieves low latency across different scenarios. Moreover, ToCC significantly reduces buffer occupancy by 4.8 times compared to existing methods, and during incast scenarios, it decreases flow completion time by up to 90%.
Chang Liu 0001, Qingyue Wang, Lu Lu 0016, Xiaoliang Wang 0001, Fu Xiao 0001, Ying Zhang 0022, Wan-Chun Dou, Guihai Chen, Chen Tian 0001
IEEE Trans. Netw.6
2025 Thunder: Minimum I/O Latency of Disaggregated Storage by Packet-Level Write-Through
abstract
The state-of-the-art storage structure relies on the NVMe devices and SmartNICs to provide high IO performance and low CPU overhead. In data centers, the existing data transmission control and storage methods are not ideal, resulting in long flow completion time, especially for small IO, which directly affects the performance of disaggregated storage systems. In this paper, we present Thunder, a disaggregated storage solution designed to minimize tail latency. Firstly, Thunder achieves the minimum I/O tail latency for disaggregated storage via packet-level write-through, and has an ingenious mechanism for precise semantic conversion from message level to packet level. It refers to the process of converting message level data into packet level data and ensuring the integrity and reliability of data transmission. This process involves steps such as message segmentation, addressing, acknowledgment, and reassembly. Secondly, we present a novel optimization approach for end-to-end and information transmission processes, aiming to address a range of issues such as user usage, congestion control, and system compatibility. Finally, we conducted both testbed and large-scale simulations to verify the performance of Thunder. The results show that Thunder reduced the average latency and tail latency by 71.6% and 59.7%, respectively compared to Gimbal and Timely. Furthermore, it effectively avoids queue head blocking and congestion diffusion in PFC, increasing throughput by 2.5X and reducing tail latency by an average of 49.7%.
Fu Xiao 0001, Weibei Fan, Xin He 0010, Junchang Wang, Xiaoliang Wang 0001, Chen Tian 0001
IEEE Trans. Netw.6
2025 Accelerating Network Features Deployment With Heterogeneous Platforms
abstract
Enhancing the networking system with appropriate functions is a longstanding goal. Unfortunately, in today’s large-scale high-speed data centers, the feature velocity of network functions is slow because it is hard to verify the function in realistic scenarios. Recent advances in programmable switching ASICs have enabled the network data plane to move beyond its traditional role of packet forwarding. However, the current compromise between performance and flexibility results in limitations such as restricted memory/computation resources and programmable models. These limitations make it challenging for programmable switches to offer more features and to be deployed in large-scale production environments. In response, we present CLIP, a framework that works in collaboration with programmable devices and commodity servers to enhance the validation and deployment velocity of features. CLIP defines a cross-platform function definition framework and provides a set of tools to reduce the complexity of manually writing cross-platform programs. We propose an automatic traffic placement and scaling mechanism to coordinate packet processing performance across heterogeneous devices. Compared with software-based Network Functions (NFs), CLIP achieves a throughput ranging from$1.36\times $to$16.06\times $under different realistic traffic loads. Through the development and deployment of three self-defined functions within a realistic testbed, we demonstrate the feasibility and efficiency of CLIP.
Xiaoliang Wang 0001, Chen Tian 0001, Yun Xiong, Sanglu Lu, Cam-Tu Nguyen
IEEE Trans. Netw.2
2024 Unison: A Parallel-Efficient and User-Transparent Network Simulation Kernel
abstract
Discrete-event simulation (DES) is a prevalent tool for evaluating network designs. Although DES offers full fidelity and generality, its slow performance limits its application. To speed up DES, many network simulators employ parallel discrete-event simulation (PDES). However, adapting existing network simulation models to PDES requires complex reconfigurations and often yields limited performance improvement. In this paper, we address this gap by proposing a parallel-efficient and user-transparent network simulation kernel, Unison, that adopts fine-grained partition and load-adaptive scheduling optimized for network scenarios. We prototype Unison based on ns-3. Existing network simulation models of ns-3 can be seamlessly transitioned to Unison. Testbed experiments on commodity servers demonstrate that Unison can achieve a 40× speedup over DES using 24 CPU cores, and a 10× speedup compared with existing PDES algorithms under the same CPU cores.
Songyuan Bai, Chen Tian 0001, Xiaoliang Wang 0001, Chang Liu 0001, Xin Jin 0008, Fu Xiao 0001, Qiao Xiang, Wan-Chun Dou, Guihai Chen
EuroSys4
2024 Sample Efficiency Matters: Training Multimodal Conversational Recommendation Systems in a Small Data Setting
abstract
With the increasing prevalence of virtual assistants, multimodal conversational recommendation systems (multimodal CRS) becomes essential for boosting customer engagement, improving conversion rates, and enhancing user satisfaction. Yet conversational samples, as training data for such a system, are difficult to obtain in large quantities, particularly in new platforms. To effectively train multimodal CRS in a small data setting, we enhance data quality to make up for the small data quantity by augmenting conversations with dialogue states. We then devise an effective dialogue state encoder to bridge the semantic gap between conversation and product representations for recommendation. To further reduce the cost of dialogue state annotation, a semi-supervised learning method is developed to effectively train the dialogue state encoder with a small set of labeled conversations. In addition, we design a correlation regularisation that leverages knowledge in the multimodal product database to help align textual and visual modalities. Experiments on the dataset MMD demonstrate the effectiveness of our method. Particularly, with only 5% of the MMD training set, our method (namely SeMANTIC) obtains better NDCG scores than those of baseline models trained on the full MMD training set.
Wenzhe Du, Xiaoliang Wang 0001, Cam-Tu Nguyen
ACM Multimedia3
2024 μMon: Empowering Microsecond-level Network Monitoring with Wavelets
abstract
Network monitoring is essential for network management and optimization. In modern data centers, fluctuations in flow rates and network congestion events (e.g., microbursts) typically manifest on a microsecond timescale. However, the time granularity of network monitoring systems has not been refined correspondingly to efficiently capture these behaviors. Attaining the monitoring granularity at the microsecond scale can greatly facilitate network performance analysis and management, but poses considerable challenges regarding memory, bandwidth, and deployment costs. We propose μMon, a novel microsecond-level network monitoring system for data centers. The key of μMon is WaveSketch, an innovative algorithm that measures and compresses flow rate curves using in-dataplane wavelet transform. WaveSketch allows for more accurate characterization of application traffic patterns and aids in profiling transport algorithms. Furthermore, by combining the fine-grained flow rate measurements with network-collected congestion information, μMon can 'replay' congestion events to analyze their cause and impact. We evaluate μMon through testbed deployment and simulations at a granularity of 8.192 μs. The evaluation results demonstrate that μMon can achieve a 90% accuracy in microsecond-level rate measurements with an average of 5 Mbps bandwidth overhead per host. Additionally, it can capture 99% heavy congestion events with 31--82 Mbps bandwidth overhead per switch.
Chengyuan Huang, Xiangyu Han, Jiaqi Zheng 0001, Xiaoliang Wang 0001, Chen Tian 0001, Wan-Chun Dou, Guihai Chen
SIGCOMM5
2024 CyberStar: Simple, Elastic and Cost-Effective Network Functions Management in Cloud Network at Scale
Bengbeng Xue, Yang Song 0031, Xiaoxin Peng, Yilong Lyu, Xiaoliang Wang 0001, Chen Tian 0001, Cam-Tu Nguyen, Biao Lyu, Rong Wen, Zhigang Zong, Shunmin Zhu
USENIX ATC7
2024 MpScope: Enabling multi-pipeline monitoring inside a switch
Chengyuan Huang, Tianfan Zhang, Li Wang 0110, Yibo Xiao, Chen Tian 0001, Xiaoliang Wang 0001, Bingheng Yan, Ahmed M. Abdelmoniem, Wan-Chun Dou, Guihai Chen
Comput. Networks7
2024 Minimizing Buffer Utilization for Lossless Inter-DC Links
abstract
RDMA over Converged Ethernet (RoCEv2) has been widely deployed to data centers (DCs) for its better compatibility with Ethernet/IP than Infiniband (IB). As cross-DC applications emerge, they also demand high throughput, low latency, and lossless network for cross-DC data transmission. However, RoCEv2’s underlying lossless mechanism Priority-based Flow Control (PFC) cannot fit into the long-haul transmission scenario and degrades the performance of RoCEv2. PFC is myopic and only considers queue length to pause upstream senders, which leads to large queueing delay. This paper proposes Bifrost, a downstream-driven lossless flow control that supports long distance cross-DC data transmission. Bifrost uses virtual incoming packets, which indicates the upper bound of in-flight packets, together with buffered packets to control the flow rate. It minimizes the buffer space requirement to one-hop bandwidth delay product (BDP) and achieves low one-way latency. Moreover, we extend Bifrost and propose BifrostX, to accommodate the multi-priority queue of the current switch implementation. BifrostX enables flow control for each queue separately while maintaining low buffer reservation, no throughput loss, and no packet loss. Real-world experiments are conducted with prototype switches and 80 kilometers cables. Evaluations demonstrate that compared to PFC, Bifrost reduces average/tail flow completion time (FCT) of inter-DC flows by up to 22.5%/42.0%, respectively. Bifrost is compatible with existing infrastructure and can support distance of thousands of kilometers.
Chengyuan Huang, Feiyang Xue, Xiaoliang Wang 0001, Tao Wu 0011, Zifa Han, Xiangyu Gong, Chen Tian 0001, Wan-Chun Dou, Guihai Chen
IEEE/ACM Trans. Netw.4
2023 MINA: Auto-scale In-network Aggregation for Machine Learning Service
Shichen Dong, Zhixiong Niu, Mingchao Zhang, Zhiying Xu, Chuntao Hu, Wei Wang 0002, Pengzhi Zhu, Qingchun Song, Peng Cheng 0005, Yongqiang Xiong, Chen Tian 0001, Cam-Tu Nguyen, Xiaoliang Wang 0001
APNet14
2023 FastWake: Revisiting Host Network Stack for Interrupt-mode RDMA
abstract
Polling and interrupt has long been a trade-off in RDMA systems. Polling has lower latency but each CPU core can only run one thread. Interrupt enables time sharing among multiple threads but has higher latency. Many applications such as databases have hundreds of threads, which is much larger than the number of cores. So, they have to use interrupt mode to share cores among threads, and the resulting RDMA latency is much higher than the hardware limits. In this paper, we analyze the root cause of high costs in RDMA interrupt delivery, and present FastWake, a practical redesign of interrupt-mode RDMA host network stack using commodity RDMA hardware, Linux OS, and unmodified applications. Our first approach to fast thread wake-up completely removes interrupts. We design a per-core dispatcher thread to poll all the completion queues of the application threads on the same core, and utilize a kernel fast path to context switch to the thread with an incoming completion event. The approach above would keep CPUs running at 100% utilization, so we design an interrupt-based approach for scenarios with power constraints. Observing that waking up a thread on the same core as the interrupt is much faster than threads on other cores, we dynamically adjust RDMA event queue mappings to improve interrupt core affinity. In addition, we revisit the kernel path of thread wake-up, and remove the overheads in virtual file system (VFS), locking, and process scheduling. Experiments show that FastWake can reduce RDMA latency by 80% on x86 and 77% on ARM at the cost of < 30% higher power utilization than traditional interrupts, and the latency is only 0.3 ∼ 0.4 μ s higher than the limits of underlying hardware. When power saving is desired, our interrupt-based approach can still reduce interrupt-mode RDMA latency by 59% on x86 and 52% on ARM.
Bojie Li, Zihao Xiang, Xiaoliang Wang 0001, Han Ruan, Jingbin Zhou, Kun Tan 0002
APNet3
2023 Dilemma of Proactive Congestion Control Protocols
abstract
Reactive congestion control (RCC) protocols have undergone decades of evolution, where senders first send data packets and then back off when congestion occurs. Recently, there has been a surge of interest in proactive congestion control (PCC) that allocates bandwidth before transmission. Despite its potential, we found that there are certain scenarios where PCC may fall short. In this paper, we aim to provide a comprehensive understanding of PCC and motivate further exploration of this area. We conduct case studies and leverage NS3 simulations to compare state-of-the-art PCC with RCCs, delving into the real dilemma of PCC.
Chen Tian 0001, Xiaoliang Wang 0001, Wan-Chun Dou, Guihai Chen
APNet3
2023 AFNFA: An Approach to Automate NCCL Configuration Exploration
abstract
With the continuously increasing scale of deep neural network models, there is a clear trend towards distributed DNN model training. State-of-the-art training frameworks support this approach using collective communication libraries such as NCCL, MPI, Gloo, and Horovod. These libraries have many parameters that can be adjusted to fit different hardware environments, and these parameters can greatly impact training performance. Therefore, careful tuning of parameters for each training environment is required. However, given the large parameter space, manual exploration can be time-consuming and laborious.
Chen Tian 0001, Xiaoliang Wang 0001, Xianping Chen
APNet4
2023 ALT: Breaking the Wall between Data Layout and Loop Optimizations for Deep Learning Compilation
abstract
Deep learning models rely on highly optimized tensor libraries for efficient inference on heterogeneous hardware. Current deep compilers typically predetermine layouts of tensors and then optimize loops of operators. However, such unidirectional and one-off workflow strictly separates graph-level optimization and operator-level optimization into different system layers, missing opportunities for unified tuning.
Zhiying Xu, Jiafan Xu, Hongding Peng, Wei Wang 0002, Xiaoliang Wang 0001, Haoran Wan, Haipeng Dai 0001, Yixu Xu, Hao Cheng 0004, Kun Wang 0005, Guihai Chen
EuroSys5
2023 Fisc: A Large-scale Cloud-native-oriented File System
Qiang Li 0045, Lulu Chen, Xiaoliang Wang 0001, Qiao Xiang, Wenhui Yao, Minfei Huang, Puyuan Yang, Shanyang Liu, Zhaosheng Zhu, Huayong Wang, Haonan Qiu, Derui Liu, Shaozong Liu, Yaohui Wu, Zhiwu Wu, Zicheng Luo, Yuchao Shao, Gexiao Tian, Zhongjie Wu, Zheng Cao 0003, Jiwu Shu, Jie Wu 0003, Jiesheng Wu
FAST3
2023 Joint Video Transcoding and Representation Selection for Edge-Assisted Multi-party Video Conferencing
Fanhao Kong, Tuo Cao, Zhuzhong Qian, Xiaoliang Wang 0001, Zhenjie Lin
ICA3PP (1)4
2023 MC-RDMA: Improving Replication Performance of RDMA-based Distributed Systems with Reliable Multicast Support
abstract
Remote Direct Memory Access has been widely adopted in distributed storage systems. However, it only supports unicast operations, which degrades the performance significantly for data replication because of bandwidth waste and CPU overhead. To address the problem, we propose MC-RDMA, a distributed and reliable multicast RDMA. It is compatible with existing unicast RDMA but supports lazy packet replication with reliable RDMA multicasting. The key idea of MC-RDMA is utilizing in-network programmable switches to build a NIC-transparent reliable multicast protocol for RDMA. MC-RDMA combines the address information of the IP and RoCEv2 into a sender-initialized multicast routing protocol. Besides, it synchronizes the hardware transmission states of multiple receivers by merging ACKs and NAKs. To verify the effectiveness of MC-RDMA, we implement it with Mellanox ConnectX-6 commodity RNICs and Intel Tofino P4 programmable switches. Experimental results show that MC-RDMA can double the sender bandwidth utilization and reduce the CPU overhead significantly compared to unicast-based RDMA replications. Moreover, it reduces the storage request latency by -30% with realistic workloads and decreases the training time by -50% in the distributed training system.
Chengyuan Huang, Yixiao Gao, Duoxing Li, Yibo Xiao, Ruyi Zhang 0005, Chen Tian 0001, Xiaoliang Wang 0001, Wan-Chun Dou, Guihai Chen, Fu Xiao 0001
ICNP8
2023 Bifrost: Extending RoCE for Long Distance Inter-DC Links
abstract
RDMA over Converged Ethernet (RoCEv2) has been widely deployed to data centers (DCs) for its better compatibility with Ethernet/IP than Infiniband (IB). As cross-DC applications emerge, they also demand high throughput, low latency, and lossless network for cross-DC data transmission. However, RoCEv2's underlying lossless mechanism Priority-based Flow Control (PFC) cannot fit into the long-haul transmission scenario and degrades the performance of RoCEv2. PFC is myopic and only considers queue length to pause upstream senders, which leads to large queueing delay. This paper proposes Bifrost, a downstream-driven lossless flow control that supports long distance cross-DC data transmission. Bifrost uses virtual incoming packets, which indicates the upper bound of in-flight packets, together with buffered packets to control the flow rate. It minimizes the buffer space requirement to one-hop bandwidth delay product (BDP) and achieves low one-way latency. Real-world experiments are conducted with prototype switches and 80 kilometers cables. Evaluations demonstrate that compared to PFC, Bifrost reduces average/tail flow completion time (FCT) of inter-DC flows by up to 22.5%/42.0%, respectively. Bifrost is compatible with existing infrastructure and can support distance of thousands of kilometers.
Feiyang Xue, Chen Tian 0001, Xiaoliang Wang 0001, Tao Wu 0011, Zifa Han, Xiangyu Gong, Wan-Chun Dou, Guihai Chen
ICNP4
2023 CLIP: Accelerating Features Deployment for Programmable Switch
abstract
Cloud network serves a large number of tenants and a variety of applications. The continuously changing demands require a programmable data plane to achieve fast feature velocity. However, the years-long release cycle of traditional function-fixed switches can not meet this requirement. Emerging programmable switches provide the flexibility of packet processing without sacrificing hardware performance. Due to the trade-off between performance and flexibility, the current programmable switches make compromises in some aspects such as limited memory/computation resources, and lack of the capacity to realize complicated computation. The programmable switches can not satisfy the demand for network services and applications in production networks. We propose a framework that leverages host servers to extend the capability of network switches quickly, accelerates new feature deployment, and verifies new ideas in production networks. Specifically, to build the unified programmable data plane, we propose essential design and implementation challenges including a programming abstraction that allows automatically and effectively deploying network functions on switch and server clusters, allocating traffic to fully utilize the server resources, and supporting flexible scaling of the system. The quick deployment of self-defined functions in a realistic system has verified the feasibility and practicality of the proposed framework.
Xiaoliang Wang 0001, Chen Tian 0001, Yun Xiong
INFOCOM2
2023 Achieving Zero-copy Serialization for Datacenter RPC
abstract
Remote Procedure Call (RPC) is widely used in distributed systems and it usually needs to serialize data before transmission. Serialization accounts for a large proportion of the overhead in RPC and becomes a bottleneck for RPC communications. Because the size of the output serialized message cannot be predicted in advance, there could be multiple memory reallocations and copies in typical serialization libraries (e.g., FlatBuffers), which dominates the overhead. We propose the novel serialization library, zFlatBuffers, to eliminate these avoidable copies during the serialization process and realize zero copy during communication. Unlike the typical serialization library, FlatBuffers, the message generated by zFlatBuffers consists of multiple non-contiguous buffers due to its zero-copy nature. Moreover, we integrate zFlatBuffers with RDMA-based RPC systems. For RDMA Unreliable Datagram, we modify the message buffer of eRPC to enable it to transmit messages composed of multiple buffers. We also build the zRPC system based on RDMA Reliable Connection, which transmits the zFlatBuffers message by the scatter/gather function. Compared to the original FlatBuffers, zFlatBuffers improves the throughput of eRPC and zRPC by 11.2%-33.7% and 5.8%-53.6%, respectively.
Tianfan Zhang, Huaping Zhou, Chengyuan Huang, Chen Tian 0001, Xiaoliang Wang 0001, Ahmed M. Abdelmoniem, Matthew Tan, Wan-Chun Dou, Guihai Chen
IPCCC6
2023 Flor: An Open High Performance RDMA Framework Over Heterogeneous RNICs
Qiang Li 0045, Yixiao Gao, Xiaoliang Wang 0001, Haonan Qiu, Yanfang Le, Derui Liu, Qiao Xiang, Bo Li 0061, Jianbo Dong, Lingbo Tang, Hongqiang Harry Liu, Shaozong Liu, Rui Miao 0001, Yaohui Wu, Zhiwu Wu, Zheng Cao 0003, Zhongjie Wu, Chen Tian 0001, Guihai Chen, Dennis Cai, Jiaji Zhu, Jiesheng Wu, Jiwu Shu
OSDI3
2022 Tuning Target Delay for RTT-based Congestion Control
abstract
The congestion control strategy plays an essential role in the high-speed datacenter network. It aims to deliver low latency, high throughput network service. RTT-based congestion control leverages advanced NIC hardware to identify accumulated queuing delay of the end-to-end path. Sender adjusts the sending rate or congestion window if the delay exceeds a predetermined value, i.e., target delay. Therefore, setting the target delay is the key for RTT-based congestion control strategies. We provide a comprehensive study of the impact of target delay on recent RTT-based congestion control strategies, and demonstrate that a fixed inappropriate target value can lead to low bandwidth utilization or high latency. We then propose a practical queuing target updating approach to solve this problem. The proposed method maintains a shared near-optimal queuing target at the receiving host. We leverage the widely supported ECN flag to estimate the empty state of switch queue instead of indicating congestion, which requires no complicated threshold configuration. We have integrated the dynamic queuing target updating approach into the state-of-the-art RTT-based congestion strategy, SWIFT, and named the design RET. Test-bed experiments and simulations in the large-scale network with synthesized traffic of real workloads show that RET can achieve up to 1.5x and 3.6x lower tail latency than SWIFT and DCQCN, respectively. This paper provides a deep understanding on tuning target delay for RTT-based congestion control algorithms in datacenter networks.
Cam-Tu Nguyen, Xiaoliang Wang 0001, Sanglu Lu
ICNP4
2022 Stabilizing and boosting I/O performance for file systems with journaling on NVMe SSD
Lin Qian, Bin Tang 0002, Xiaoliang Wang 0001, Sanglu Lu
Sci. China Inf. Sci.5
2022 FastCache: A write-optimized edge storage system via concurrent merging cache for IoT applications
Lin Qian, Zhihao Qu, Miao Cai 0001, Xiaoliang Wang 0001, Weiguo Duan
J. Syst. Archit.5
2021 FastCache: A Client-Side Cache with Variable-Position Merging Schema in Network Storage System
Lin Qian, Xiaoliang Wang 0001, Zhihao Qu, Weiguo Duan
ICA3PP (2)3
2021 Maximizing the Benefit of RDMA at End Hosts
abstract
RDMA is increasingly deployed in data center to meet the demands of ultra-low latency, high throughput and low CPU overhead. However, it is not easy to migrate existing applications from the TCP/IP stack to the RDMA. The developers usually need to carefully select communication primitives and manually tune the parameters for each single-purpose system. After operating the high-speed RDMA network, we identify multiple hidden costs which may cause degraded and/or unpredictable performance of RDMA-based applications. We demonstrate these hidden costs including the combination of complicated parameter settings, scalability of Reliable Connections, two-sided memory management and page alignment, resource contention among diverse traffics, etc. Furthermore, to address these problems, we introduce Nem, a suite that allows developers to maximize the benefit of RDMA by i) eliminating the resource contention at NIC cache through asynchronous resource sharing; ii) introducing hybrid page management based on messages sizes; iii) isolating flows of different traffic classes based hardware features. We implement the prototype of Nem and verify its effectiveness by rebuilding the RPC message service, which demonstrates the high throughput for large messages, low latency for small messages without compromising the low CPU utilization and good scalability performance for a large number of active connections.
Xiaoliang Wang 0001, Hexiang Song, Cam-Tu Nguyen, Dongxu Cheng, Tiancheng Jin
INFOCOM1
2021 Accessing Cloud with Disaggregated Software-Defined Router
Xiaoliang Wang 0001, Yuanwei Lu, Yanbo Yu, Shengli Zheng, Youjian Zhao
NSDI2
2021 ACC: automatic ECN tuning for high-speed datacenter networks
abstract
For the widely deployed ECN-based congestion control schemes, the marking threshold is the key to deliver high bandwidth and low latency. However, due to traffic dynamics in the high-speed production networks, it is difficult to maintain persistent performance by using the static ECN setting. To meet the operational challenge, in this paper we report the design and implementation of an automatic run-time optimization scheme, ACC, which leverages the multi-agent reinforcement learning technique to dynamically adjust the marking threshold at each switch. The proposed approach works in a distributed fashion and combines offline and online training to adapt to dynamic traffic patterns. It can be easily deployed based on the common features supported by major commodity switching chips. Both testbed experiments and large-scale simulations have shown that ACC achieves low flow completion time (FCT) for both mice flows and elephant flows at line-rate. Under heterogeneous production environments with 300 machines, compared with the well-tuned static ECN settings, ACC achieves up to 20\% improvement on IOPS and 30\% lower FCT for storage service. ACC has been applied in high-speed datacenter networks and significantly simplifies the network operations.
Xiaoliang Wang 0001, Yinben Xia, Derui Liu, Weishan Deng
SIGCOMM2
2021 Gray Failures Detection for Shared Bicycles
Hangfan Zhang, Mingchao Zhang, Cam-Tu Nguyen, Sheng Zhang 0001, Xiaoliang Wang 0001
WASA (1)6
2021 SAKE: Estimating Katz Centrality Based on Sampling for Large-Scale Social Networks
abstract
Katz centrality is a fundamental concept to measure the influence of a vertex in a social network. However, existing approaches to calculating Katz centrality in a large-scale network are unpractical and computationally expensive. In this article, we propose a novel method to estimate Katz centrality based on graph sampling techniques, which object to achieve comparable estimation accuracy of the state-of-the-arts with much lower computational complexity. Specifically, we develop a Horvitz–Thompson estimate for Katz centrality by using a multi-round sampling approach and deriving an unbiased mean value estimator. We further propose SAKE , a S ampling-based A lgorithm for fast K atz centrality E stimation. We prove that the estimator calculated by SAKE is probabilistically guaranteed to be within an additive error from the exact value. Extensive evaluation experiments based on four real-world networks show that the proposed algorithm can estimate Katz centralities for partial vertices with low sampling rate, low computation time, and it works well in identifying high influence vertices in social networks.
Mingkai Lin, Lynda Jiwen Song, Cam-Tu Nguyen, Xiaoliang Wang 0001, Sanglu Lu
ACM Trans. Knowl. Discov. Data5
2020 RouteStitch: Control Traffic Minimization in SDN by Stitching Routes
abstract
Software Defined Networking (SDN) is beneficial to many applications, such as intra-datacenter communication, inter-datacenter transportation, etc., due to its centralized control. However, this centralized control frequently makes the controller a bottleneck, due to the large amount of interactions between the controller and switches. In this paper, we characterize such interactions as control traffic, and propose RouteStitch to minimize such kind of traffic. RouteStitch exploits existing route entries in switches to build new paths. To this end, RouteStitch first builds a graph model to describe existing route entries. Then, on such a model, a novel minimum color-alternation routing problem is defined to minimize control traffic, after which an optimal algorithm is proposed on a fixed routing path. For general paths, an O(log2L)-competitive online algorithm is designed to build new paths in an online manner that preserves fundamental property of switch Ternary Content Addressable Memory (TCAM) capacity and allowed maximum hop length L. Extensive simulation results based on realistic topology show that RouteStitch has good performance in terms of reducing control traffic, by 40%.
An Xie, Huawei Huang, Xiaoliang Wang 0001, Zhuzhong Qian, Sanglu Lu
ICC3
2020 Resource-Efficient and Convergence-Preserving Online Participant Selection in Federated Learning
abstract
Federated learning achieves the privacy-preserving training of models on mobile devices by iteratively aggregating model updates instead of raw training data to the server. Since excessive training iterations and model transferences incur heavy usage of computation and communication resources, selecting appropriate devices and excluding unnecessary model updates can help save the resource usage. We formulate an online time-varying non-linear integer program to minimize the cumulative resource usage over time while achieving the desired long-term convergence of the model being trained. We design an online learning algorithm to make fractional control decisions based on both previous system dynamics and previous training results, and also design an online randomized rounding algorithm to convert the fractional decisions into integers without violating any constraints. We rigorously prove that our online approach only incurs sub-linear dynamic regret for the optimality loss and sub-linear dynamic fit for the long-term convergence violation. We conduct extensive trace-driven evaluations and confirm the empirical superiority of our approach over alternative algorithms in terms of up to 27% reduction on the resource usage while sacrificing only 4% reduction on accuracy.
Yibo Jin 0001, Lei Jiao 0002, Zhuzhong Qian, Sheng Zhang 0001, Sanglu Lu, Xiaoliang Wang 0001
ICDCS6
2020 Provisioning Edge Inference as a Service via Online Learning
abstract
Provisioning machine learning inference as a service at the mobile network edge for distributed users in an online setting faces multiple challenges, including the accuracy-resource trade-off for model selection, the time-coupled decision for model distribution, and the unpredictable user inference workload. To overcome such challenges, we firstly model an online time-varying non-linear integer program of maximizing the overall service's inference accuracy through dynamic model instance selection, delivery and workload distribution. Afterwards, we design an online learning algorithm to make fractional control decisions, which alternates between minimizing an outer problem and maximizing an inner problem of an equivalent convex-concave formulation by only taking previously observable inputs. We further design a randomized rounding algorithm to convert the fractional decisions into integers. We rigorously prove that our approach only incurs sub-linear dynamic regret for the optimality loss and sub-linear dynamic fit for the long-term constraints violation. Finally, we conduct extensive evaluations with real- world data and confirm the empirical superiority of our approach over state-of-the-art algorithms in terms of up to 30% reduction on accuracy loss and 34% reduction on constraints violation.
Yibo Jin 0001, Lei Jiao 0002, Zhuzhong Qian, Sheng Zhang 0001, Ning Chen 0010, Sanglu Lu, Xiaoliang Wang 0001
SECON7
2020 App trajectory recognition over encrypted internet traffic based on deep neural network
Xiaoliang Wang 0001, Cam-Tu Nguyen, Sanglu Lu
Comput. Networks3
2020 Online VNF chain deployment on resource-limited edges by exploiting peer edge devices
An Xie, Huawei Huang, Xiaoliang Wang 0001, Zhuzhong Qian, Sanglu Lu
Comput. Networks3
2020 Construction of Subexponential-Size Optical Priority Queues With Switches and Fiber Delay Lines
abstract
All-optical switching has been considered as a natural choice to keep pace with growing fiber link capacity. One key research issue of all-optical switching is the design of optical buffers for packet contention resolution. One of the most general buffering schemes is optical priority queue, where every packet is associated with a unique priority upon its arrival and departs the queue in order of priority, and the packet with the lowest priority is always dropped when a new packet arrives but the buffer is full. In this paper, we focus on the feedback construction of an optical priority queue with a single (M + 2) × (M + 2) optical crossbar Switch and M fiber Delay Lines (SDL) connecting M inputs and M outputs of the switch. We propose a novel construction of an optical priority queue with buffer 2Θ(√M), which improves substantially over all previous constructions that only have buffers of O(Mc) size for constant integer c. The key ideas behind our construction include (i) the use of first in first out multiplexers, which admit efficient SDL constructions, for feeding back packets to the switch instead of fiber delay lines, and (ii) the use of a routing policy that is similar to self-routing, where each packet entering the switch is routed to some multiplexer mainly determined by the current ranking of its priority.
Bin Tang 0002, Xiaoliang Wang 0001, Cam-Tu Nguyen, Sanglu Lu
IEEE/ACM Trans. Netw.2
2019 Sampling Based Katz Centrality Estimation for Large-Scale Social Networks
Mingkai Lin, Cam-Tu Nguyen, Xiaoliang Wang 0001, Sanglu Lu
ICA3PP (2)4
2019 ReLeS: A Neural Adaptive Multipath Scheduler based on Deep Reinforcement Learning
abstract
The Multipath TCP (MPTCP) protocol, featured by its ability of capacity aggregation across multiple links and connectivity maintenance against single-path failure, has been attracting increasing attention from the industry and academy. Multipath packet scheduling is a unique and fundamental mechanism for the design and implementation of MPTCP, which is responsible for distributing the traffic over multiple subflows. The existing multipath schedulers are facing the challenges of network heterogeneities, comprehensive QoS goals, and dynamic environments, etc. To address these challenges, we propose ReLeS, a Reinforcement Learning based Scheduler for MPTCP. ReLeS uses modern deep reinforcement learning (DRL) techniques to learn a neural network to generate the control policy for packet scheduling. It adopts a comprehensive reward function that takes diverse QoS characteristics into consideration to optimize packet scheduling. To support real-time scheduling, we propose an asynchronous training algorithm that enables parallel execution of packet scheduling, data collecting, and neural network training. We implement ReLeS in the Linux kernel and evaluate it over both emulated and real network conditions. Extensive experiments show that ReLeS significantly outperforms the state-of-the-art schedulers.
Shaohua Gao, Xiaoliang Wang 0001
INFOCOM4
2019 ActiveTracker: Uncovering the Trajectory of App Activities over Encrypted Internet Traffic Streams
abstract
Despite the increasing popularity of mobile applications and the widespread adoption of encryption techniques, mobile devices are still susceptible to security and privacy risks. In this paper, we propose ActiveTracker, a new type of sniffing attack that can reveal the fine-grained trajectory of user’s mobile app usage from a sniffed encrypted Internet traffic stream. It firstly adopts a sliding window based approach to divide the encrypted traffic stream into a sequence of segments corresponding to different app activities. Then each traffic segment is represented by a normalized temporal-spacial traffic matrix and a traffic spectrum vector. Based on the normalized representation, a deep neural network (DNN) classification algorithm is developed to recognize the crucial activities conducted with different apps by the user. We show by extensive experiments on real-world app usage traffic collected from volunteers that the proposed approach achieves up to 78.5% accuracy in recognizing app trajectory over encrypted traffic streams.
Xiaoliang Wang 0001, Cam-Tu Nguyen, Sanglu Lu
SECON3
2019 Dual: Deploy stateful virtual network function chains by jointly allocating data-control traffic
An Xie, Huawei Huang, Xiaoliang Wang 0001, Song Guo 0001, Zhuzhong Qian, Sanglu Lu
Comput. Networks3
2019 SmartCC: A Reinforcement Learning Approach for Multipath TCP Congestion Control in Heterogeneous Networks
abstract
The Multipath TCP (MPTCP) protocol has been standardized by the IETF as an extension of conventional TCP, which enables multi-homed devices to establish multiple paths for simultaneous data transmission. Congestion control is a fundamental mechanism for the design and implementation of MPTCP. Due to the diverse QoS characteristics of heterogeneous links, existing multipath congestion control mechanisms suffer from a number of performance problems such as bufferbloat, suboptimal bandwidth usage, etc. In this paper, we propose a learning-based multipath congestion control approach called SmartCC to deal with the diversities of multiple communication path in heterogeneous networks. SmartCC adopts an asynchronous reinforcement learning framework to learn a set of congestion rules, which allows the sender to observe the environment and take actions to adjust the subflows' congestion windows adaptively to fit different network situations. To deal with the problem of infinite states in high-dimensional space, we propose a hierarchical tile coding algorithm for state aggregation and a function estimation approach for Q-learning, which can derive the optimal policy efficiently. Due to the asynchronous design of SmartCC, the processes of model training and execution are decoupled, and the learning process will not introduce extra delay and overhead on the decision making process in MPTCP congestion control. We conduct extensive experiments for performance evaluation, which show that SmartCC improves the aggregate throughput significantly and outperforms the state-of-the-art mechanisms on a variety of performance metrics.
Shaohua Gao, Chaojing Xue, Xiaoliang Wang 0001, Sanglu Lu
IEEE J. Sel. Areas Commun.5
2019 Semi-Clairvoyant Scheduling in Data Analytics Systems
abstract
Popular data analytics systems including Apache Hadoop, Dryad, and Apache Spark abstract jobs as directed acyclic graphs (DAGs). Speeding up completions for DAG jobs matter in practice in order to support real-time decisions. State-of-the-art works propose clairvoyant schedulers to optimize these goals, however, they assume complete job information as a prior knowledge which includes the precise DAG structure, and fine-grained resource requirement and duration time of each task. This assumption limits their applicability. In this paper, to be more practical, we relax the complete prior knowledge assumption and rely solely on partial prior information, based on which, we design a semi-clairvoyant task scheduler Cobra operating within each job. When managing resources for a job, Cobra adaptively adjusts its resource desires in a multiplicative-increase multiplicative-decrease manner on the basis of the its current resource utilization and the presence of current waiting tasks. When assigning tasks to run on the allocated resources, Cobra strives to satisfy task locality preference by tolerating each task waiting for some time that is bounded by a parameterized threshold. Even with the partial prior job information, when a set of jobs in which each employing Cobra as its task scheduler, run on a cluster that employs the fair job scheduler, we theoretically prove the produced makespan and average job response time are O(1)-competitive in different settings. We implement our design in Spark on YARN system, and use experiments from both real deployments and simulations on Google's trace to verify the performance promotion and sensitivity of Cobra.
Xiaoda Zhang, Zhuzhong Qian, Sheng Zhang 0001, Xiangbo Li, Xiaoliang Wang 0001, Sanglu Lu
IEEE Trans. Computers5
2018 Toward Effective and Fair RDMA Resource Sharing
abstract
Remote Direct Memory Access (RDMA) technique allows the messaging service that directly access the memory on remote machines, which provides low CPU overhead, low latency, and high throughput network transmission. On the other hand, however, due to the limited cache space in RDMA NIC (RNIC), it is still challenging to achieve effective and fair resource sharing across different applications. To address this problem, we present a scalable RDMA as a service to manage resource and deliver fair scheduling to applications' requests. We study the thread contention and preemptive schedule issues at end-hosts, and report the corresponding performance degradation through experiments. Then, we introduce Avatar, a model to manage memory and Queue Pairs (QPs) resource for a large number of connections, which eliminates the lock contention and provides fair data scheduling for applications with different priorities. Finally, we implement Avatar and demonstrate that Avatar can support a thousand of connections, improve the fairness and reduce the requests completion time up to 50% in comparison with the native RDMA.
Haonan Qiu, Xiaoliang Wang 0001, Tianchen Jin, Zhuzhong Qian, Bin Tang 0002, Sanglu Lu
APNet2
2018 UKSM: Swift Memory Deduplication via Hierarchical and Adaptive Memory Region Distilling
Nai Xia, Chen Tian 0001, Yan Luo 0001, Hang Liu 0001, Xiaoliang Wang 0001
FAST5
2018 Far-Sighted Multi-Stage Awasre Coflow Scheduling
abstract
In data center networks (DCN), large scale flows produced by parallel computing frameworks form many coflows semantically. Most inter-coflow schedulers only focus on the remaining data of coflows and attempt to mimic Shortest Job First (SJF). However, a coflow may consist of multiple stages. In this paper, we consider the Multi-stage Inter-Coflow Scheduling problem and try to give an efficient online scheduling scheme. We first explore a short-sighted algorithm with the greedy strategy. This gives us an insight into utilizing the network resources. Based on that, we propose a far-sighted heuristic, which schedules sub-coflows to occupy network bandwidth in turn. Through simulations in various network environments, we show that, compared to a state-of-the-art scheduler - Varys, a multi-stage aware scheduler can reduce the coflow completion time by up to \pmb4.81× even though it is short-sighted. Moreover, the far-sighted scheduler can improve the performance by nearly \pmb7.95 × reduction.
Shuai Zhang 0058, Sheng Zhang 0001, Xiaoda Zhang, Zhuzhong Qian, Mingjun Xiao, Jie Wu 0001, Jidong Ge, Xiaoliang Wang 0001
GLOBECOM8
2018 Optimizing User Experience through Implicit Content-aware Network Service in the Home Environment
abstract
There has always been a gap between Internet Service Providers (ISPs) and end users when considering the performance of network-based application. On one hand, ISPs keep raising the investment on infrastructures to speed up the data transportation. On the other hand, users are not satisfied with the perceived quality of experience (QoE). This happens mainly due to the inflexible network flow management, where only the function of rate limiting is provided for home users in the shared network environment. In this paper, we focus on the optimization of users experience by customizing bandwidth allocation for user specified preferences while maintaining high bandwidth utilization. We introduce implicit content-aware bandwidth allocation to minimize the involvement of users on complicated network setting. By leveraging the technique of software-defined networking (SDN), a prototype of content-aware traffic scheduling, Conan, is developed to verify the effectiveness of our design. Experiments show that Conan can reduce the average task completion time of interactive applications by 30-40%. During heavy traffic load, Conan can ensure stable bandwidth for each video streaming flow and greatly reduce the average stall duration.
Haixiang Yang, Xiaoliang Wang 0001, Cam-Tu Nguyen, Sanglu Lu
GROUP2
2018 Nem: Toward Fine-grained Load Balancing through RNIC EC Offloading
abstract
Modern datacenter networks employ Load-balancing (LB) in the large-scale multi-tier topology to ensure high network utilization as well as low flow completion time. This paper presents the design and evaluation of Nem, a robust Erasure Coding (EC) based load balancing scheme at end-host to spread data across multiple paths. Our design is based on two key insights. First, both theory and implementation have shown that redundancy is a powerful technique to reduce latency in networked system. Second, the commercial RDMA network interface card supports EC offload which can dramatically reduce the CPU consumption. Nem is an optimal user-level LB design, which leveraging redundant fine-grained data blocks and high speed lossless RDMA network to realize effective load balancing transmission. Evaluation over many workloads shows that Nem is adaptive to the asymmetric networks, and achieves better performance compared to the state-of-art host-based load balancing mechanism.
Xiaoliang Wang 0001, Cam-Tu Nguyen, Zhuzhong Qian, Bin Tang 0002, Sanglu Lu
HPSR1
2018 ran-GJS: Orchestrating Data Analytics for Heterogeneous Geo-distributed Edges
abstract
Many organizations and companies have deployed not only datacenters but also large number of geo-distributed heterogeneous edges to provide fast data analytics services. Since large volume of data transmission across WAN can be costly, existing works mainly focus on pre-processing data in-place to avoid transmission. However, the heterogeneity of edges on either local computing capacity or network bandwidth limits the efficient use on scarce resource, which may result in long task completion time. To cope with dynamic demands on scarce resource, we take the heterogeneity of both computing capacity and network bandwidth of geo-distributed edges into consideration when assigning data analytical tasks and their associated data between the central datacenter and edges such that the overall latency can be reduced. We formulate the geo-distributed data-task joint scheduling problem (GJS), show its NP-hardness, and propose a near-optimal randomized scheduling algorithm (ran-GJS). ran-GJS can be proved concentrated around its optimum value with high probability, i.e., 1--O(e--t2) where t is the concentration bound by using Martingale Analysis. The experimental results obtained form both extensive simulations and Yarn-based prototype show that ran-GJS significantly speeds up the geo-distributed analytics with a gain on average completion time of at least 28% over state-of-the-art baseline algorithms.
Yibo Jin 0001, Zhuzhong Qian, Song Guo 0001, Sheng Zhang 0001, Xiaoliang Wang 0001, Sanglu Lu
ICPP5
2018 COBRA: Toward Provably Efficient Semi-Clairvoyant Scheduling in Data Analytics Systems
abstract
Typical data analytics systems abstract jobs as directed acyclic graphs (DAGs). It is crucial to maximize throughput and speedup completions for DAG jobs in practice. Existing works propose clairvoyant schedulers optimizing these goals, however, they assume complete job information as a prior knowledge which limits their applicability. Instead, we remove the complete prior knowledge assumption and rely solely on a partial prior information, which is more practical. And we design a semi-clairvoyant task scheduler Cobra working within each job. Cobra adaptively adjusts its resource desires in a multiplicative-increase multiplicative-decrease (MIMD) manner according to nearly past resource utilizations and the current waiting tasks. On the other hand, Cobra seeks to satisfy task locality preferences by allowing each task to wait for some time that is bounded by a parameterized threshold. Surprisingly, even with the partial prior job information, we theoretically prove, Cobra, when working with the widely used fair job scheduler, is O(1)-competitive with respect to both makespan and average job response time. We experimentally validate that the performance promotion of Cobra in both real system deployment and trace-driven simulations.
Xiaoda Zhang, Zhuzhong Qian, Sheng Zhang 0001, Xiangbo Li, Xiaoliang Wang 0001, Sanglu Lu
INFOCOM5
2018 Modeling Geographically Correlated Failures to Assess Network Vulnerability
abstract
Current communication networks are facing more and more threats from large-scale regional damages, such as natural disasters (e.g., earthquake or tornado) and physical attacks (e.g., electromagnetic pulse attack or dragging anchors). Recently, several region failure models have been proposed to evaluate the impact of such geographically correlated failures on communication networks. These works mainly adopt a kind of “deterministic” models, where network components within the affected area would fail simultaneously. Such failure models simplify the analysis but may fail to reflect some important behaviors of attacks and thus cause significant over- or under-estimation of region failures. To emulate the impact of realistic catastrophe events such as earthquake and tornado, this paper introduces two probabilistic failure models: 1) a concentric circle model and 2) a line segment model. In the probabilistic models, both the location and effect of the damage are treated as random events. The failure may randomly incur on the entire network plane, and the failure probability of a device depends on many factors, e.g., link length and distance to the damage center. We develop an efficient grid partition-based scheme to estimate the network vulnerability. Based on the grid partition scheme, we further develop a sampling scheme to significantly reduce the computation cost. The probability model helps us more deeply understand the network behaviors under region failure and facilitates the design and maintenance of future highly survivable mission critical networks.
Xiaoliang Wang 0001, Sanglu Lu
IEEE Trans. Commun.1
2018 FUSO: Fast Multi-Path Loss Recovery for Data Center Networks
Guo Chen 0001, Yuanwei Lu, Yuan Meng 0002, Bojie Li, Kun Tan 0002, Dan Pei, Peng Cheng 0005, Layong Luo, Yongqiang Xiong, Xiaoliang Wang 0001, Youjian Zhao
IEEE/ACM Trans. Netw.10
2017 Ambula: Build Communication Lifeline of Corporations During Emergency
abstract
Many corporations rely on Internet service provider (ISP) network to provide reliable communication services. However, the current communication networks are vulnerable to disruptive events, such as natural disaster or power outage. Such disastrous events may destroy multiple network facilities in a specific region and result in a long term recovery of ISP networks. The disconnected communication will lead to enormous economic loss even if corporation's infrastructure is not directly destroyed during the disaster. Therefore, corporations need a self-rescue mechanism to actively respond to the emergency instead of simply relying on the ISP. This paper proposes Ambula, an easy-to-deploy platform to realize fast congestion-aware recovery for corporation's communication lifeline. Our platform leverages current widely-deployed public cloud services to build a scalable peer-to-peer overlay routing system. By so doing, the corporation is capable of controlling the packets forwarding path to bypass the affected region and congested routes. To this end, Ambula first carefully selects a small set of virtual machines (VMs) from geographically distributed public clouds, and then apply the self-developed congestion-aware routing protocol to achieve automatic and fast routing recovery. Simulations on both random generated and real network topologies show that the high recovery ratio of 80% can be achieved. The congestion avoidance algorithm can significantly reduce the impact of congestion. Our prototype on Emulab shows it can recover within hundreds of milliseconds. To the best of our knowledge, no effective disaster recovery mechanism currently exists for corporations during emergency. Ambula will facilitate the business continuity management of corporations in present of hazard events.
An Xie, Xiao Zhang 0015, Xiaoliang Wang 0001, Zhuzhong Qian, Sanglu Lu
ICPADS3
2017 A Virtual Middleboxes Network Placement Algorithm in Multi-tenant Datacenter Networks
abstract
Hardware middleboxes are widely used in current cloud datacenter to provide network functions such as firewalls, intrusion detection system, load balancers, etc. Unfortunately, they are expensive and unable to offer customized functions for individual tenant. To overcome this issue, there is an increasing interest in deploying software middleboxes to enable flexible security, network access functionality. This paper addresses the software middleboxes placement problem with minimum bandwidth guarantee. We first specify the model of tenants' requirement that specifies the need for virtual machines of application and middleboxes, as well as communication traffic. A virtual middlebox placement algorithm called MISSILE is then proposed to offer predictable network performance for each accepted tenant, and minimize datacenter bandwidth utilization. Extensive simulation results based on current large-scale datacenter networks verify that MISSILE is effective and provides network performance guarantee for tenants.
Xiaoliang Wang 0001, Cam-Tu Nguyen, Jian Wang 0038, Zhuzhong Qian, Sanglu Lu
ICPADS2
2017 AGRA: An Analysis-Generation-Ranking Framework for Automatic Abbreviation from Paper Titles
abstract
People sometimes choose word-like abbreviations to refer to items with a long description. These abbreviations usually come from the descriptive text of the item and are easy to remember and pronounce, while preserving the key idea of the item. Coming up with a nice abbreviation is not an easy job, even for human. Previous assistant naming systems compose names by applying hand-written rules, which may not perform well. In this paper, we propose to view the naming task as an artificial intelligence problem and create a data set in the domain of academic naming. To generate more delicate names, we propose a three-step framework, including description analysis, candidate generation and abbreviation ranking, each of which is parameterized and optimizable. We conduct experiments to compare different settings of our framework with several analysis approaches from different perspectives. Compared to online or baseline systems, our framework could achieve the best results.
Shujian Huang, Cam-Tu Nguyen, Xiaoliang Wang 0001, Xinyu Dai, Jiajun Chen 0001, Yang Yu 0001
IJCAI5
2017 One more queue is enough: Minimizing flow completion time with explicit priority notification
abstract
Ideally, minimizing the flow completion time (FCT) requires millions of priorities supported by the underlying network so that each flow has its unique priority. However, in production datacenters, the available switch priority queues for flow scheduling are very limited (merely 2 or 3). This practical constraint seriously degrades the performance of previous approaches. In this paper, we introduce Explicit Priority Notification (EPN), a novel scheduling mechanism which emulates fine-grained priorities (i.e., desired priorities or DP) using only two switch priority queues. EPN can support various flow scheduling disciplines with or without flow size information. We have implemented EPN on commodity switches and evaluated its performance with both testbed experiments and extensive simulations. Our results show that, with flow size information, EPN achieves comparable FCT as pFabric that requires clean-slate switch hardware. And EPN also outperforms TCP by up to 60.5% if it bins the traffic into two priority queues according to flow size. In information-agnostic setting, EPN outperforms PIAS with two priority queues by up to 37.7%. To the best of our knowledge, EPN is the first system that provides millions of priorities for flow scheduling with commodity switches.
Yuanwei Lu, Guo Chen 0001, Larry Luo, Kun Tan 0002, Yongqiang Xiong, Xiaoliang Wang 0001, Enhong Chen
INFOCOM6
2017 SilentTalk: Lip reading through ultrasonic sensing on mobile phones
abstract
The recently enhanced computing capability and rich sensing functionality on mobile devices lead to the ubiquitous application of speech recognition. Traditional speech recognition records acoustic signals or visual images to interpret speech. However, the acoustic based scheme has many drawbacks. It is easily affected by the environmental noise when users are in the factory or market, and can not be used in a place where people need to be quite such as library. Specifically the current design is not suitable for people with speaking or hearing difficulties. Unfortunately, the visual-based approach is sensitive to fight conditions which shows poor performance in the dark area. As a result, it is necessary to provide an new human-computer interaction channel to assist speech recognition. This paper presents SilentTalk, a non-invasive lip reading system based on ultrasonic Doppler effect The main idea is to generate ultrasonic signals from a mobile phone, then capture the reflections and analyze the fine-grained frequency shift caused by mouth movements. A Frequency Shift Detection Model (FSDM) is proposed to quantify the correlation between frequency variations and mouth movements that form different syllables. SilentTalk then applies a Continuous Lip Reading Model (CLRM) on top of FSDM to realize continuous lip reading. Based on Markov assumption, CLRM effectively combines pronunciation rules and context knowledge to connect isolated syllables to words and sentences. Experiments show that SilentTalk can identify 12 basic mouth motions up to 95.4% accuracy in English. The system can also recognize short sentences up to six words with an average accuracy of 74.8%.
Jiayao Tan, Cam-Tu Nguyen, Xiaoliang Wang 0001
INFOCOM3
2017 Rethinking transfer optimization in a datacenter: Integrating load balancing with multipath flow control
abstract
The various flows in production datacenters usually can be classified into two types: bandwidth-hungry and delay-sensitive. To improve their performance, datacenter networks require effective load balancing and flow control protocols, respectively. However, as the two techniques are typically employed separately in current datacenters, they are unable to optimize the network in a coordinated way. In this work, we argue that the adaptive routing, in load balancing sense, and the flow control, in congestion control sense, could be tightly coupled at the transport layer to handle the complex datacenter traffic. We design OmniFlow, a novel transfer protocol which aims to achieve a proper balance between throughput and latency in a datacenter. Firstly, it can simultaneously and precisely measure the queueing latencies on multiple paths between two hosts, which enables it to have more visibility of the path congestion and have better control of the transmission states. Secondly, OmniFlow adaptively integrates the load balancing and flow control modules and shares the same congestion metrics (i.e. queueing latencies) between them. Based on different network conditions, it either dynamically reroutes flows to utilize the bisection bandwidth or proactively adjusts flow rates to bound queueing occupancies. The results of extensive experiments show that OmniFlow can provide both low average and tail latency for small flows without sacrificing the throughput of elephant flows.
Zhuzhong Qian, Kaiyuan Wen, Sheng Zhang 0001, Xiaoliang Wang 0001, Sanglu Lu
IWQoS4
2017 Predicting Happiness State Based on Emotion Representative Mining in Online Social Networks
Xiao Zhang 0015, Hong Huang 0001, Cam-Tu Nguyen, Xu Chen 0004, Xiaoliang Wang 0001, Sanglu Lu
PAKDD (1)6
2016 Constructing sub-exponentially large optical priority queues with switches and fiber delay lines
abstract
Optical switching has been considered as a natural choice to keep pace with growing fiber link capacity. One key research issue of all-optical switching is the design of optical queues by using optical crossbar switches and fiber delay lines (SDLs). In this paper, we focus on the construction of an optical priority queue with a single (M+2)×(M+2) crossbar switch and M fiber delay lines, and evaluate it in terms of the buffer size of the priority queue. Currently, the best known upper bound of the buffer size is O(2M), while existing methods can only construct a priority queue with buffer O(M3). In this paper, we make a great step towards closing the above huge gap. We propose a very efficient construction of priority queues with buffer 2Θ(√M). We use 4-to-1 multiplexers with different buffer sizes, which can be constructed efficiently with SDL, as intermediate building blocks to simplify the design. The key idea in our construction is to route each packet entering the switch to some group of four 4-to-1 multiplexers according to its current priority, which is shown to be collision-free.
Bin Tang 0002, Xiaoliang Wang 0001, Cam-Tu Nguyen, Sanglu Lu
ISIT2
2016 An Efficient Walking Safety Service for Distracted Mobile Users
abstract
There is a growing number of incidents related to using cell phones while walking on the street. To address this issue, this paper proposes a new system based on tactile paving detection on the sidewalk to alert distracted mobile users to avoid traffic hazard. Our system (namely Inspector) is deployed as an application on an off-the-shelf normal mobile phone equipped with back camera. Inspector plays as a third eye to alert users when they step out of the safe zones where no tactile paving is detected. In order to obtain reliable yet effective decision results, we exploit a lightweight image processing method, simple classifiers along with a smart sampling strategy. The main idea is that the application capture more images of the surrounding environments when we doubt that the result from one image is not sufficient. The sampling interval is adjusted dynamically so that we can save the energy and thus extend the working time of the system. Real-scenario tests show that Inspector can detect whether a mobile user is walking along a blind sidewalk with an accuracy of 92.72%, 98.78%, and 99.44% when the detection algorithm sample 2 times, 3 times, and 4 times continuous detection respectively. The reaction time, which is measured by the difference between the alert time and the time when a user steps out of a safety zone, is 0.52 seconds early with a sampling interval of 2 seconds.
Maozhi Tang, Cam-Tu Nguyen, Xiaoliang Wang 0001, Sanglu Lu
MASS3
2016 Riding quality evaluation through mobile crowd sensing
abstract
Public transport plays an importation role in our daily life. The information related to passengers satisfaction is very beneficial for optimizing the transportation service. This paper investigates an application of mobile crowd sensing to detect and analyze the riding quality of public transport vehicles. The lightweight system leverages sensors equipped on participants' smartphones to collect surrounding information. By analyzing the uploaded data at a server, we are able to estimate both aggressive driving behaviors and environment contexts. Series of data processing methods are exploited to overcome the affection of body movement and road condition, and crowd sourcing is applied to improve the robustness of the results. We have tested this system in 3 different transportation in 3 cities. The results indicate that the system can provide sufficient accuracy (up to 91% with 7 phones) to identify dozens of riding-comfort metrics.
Senyuan Tan, Xiaoliang Wang 0001, Guido Maier
PerCom2
2016 Conan: Content-aware Access Network Flow Scheduling to Improve QoE of Home Users
abstract
There has always been a gap of perception between Internet Service Providers (ISPs) and their customers when considering the performance of network service. On one hand, ISPs invest to increase downstream speed of access network infrastructure. On the other hand, users cannot achieve perceived quality of experience (QoE). This paper addresses this problem by introducing a system, Conan, which enables content-aware flow scheduling to improve the QoE of users. Conan exploits to satisfy users' requirements in the access network (LAN), which is the performance bottleneck actually. By leveraging the technique of software defined networking (SDN), Conan are able to specify the expected network capacity for different applications. Automatic application identification is deployed at home gateway to improve the scalability, and flexible bandwidth allocation is realized at LAN for specified applications. Using video streaming service optimization as an example, we demonstrate that our system can automatically allocate bandwidth for video flows.
Haixiang Yang, Xiaoliang Wang 0001, Cam-Tu Nguyen, Sanglu Lu
SIGCOMM2
2016 Fast and Cautious: Leveraging Multi-path Diversity for Transport Loss Recovery in Data Centers
Guo Chen 0001, Yuanwei Lu, Yuan Meng 0002, Bojie Li, Kun Tan 0002, Dan Pei, Peng Cheng 0005, Layong Luo, Yongqiang Xiong, Xiaoliang Wang 0001, Youjian Zhao
USENIX ATC10
2014 Labeling Complicated Objects: Multi-View Multi-Instance Multi-Label Learning
abstract
Multi-Instance Multi-Label (MIML) is a learning framework where an example is associated with multiple labels and represented by a set of feature vectors (multiple instances). In the formalization of MIML learning, instances come from a single source (single view). To leverage multiple information sources (multi-view), we develop a multi-view MIML framework based on hierarchical Bayesian Network, and derive an effective learning algorithm based on variational inference. The model can naturally deal with examples in which some views could be absent (partial examples). On multi-view datasets, it is shown that our method is better than other multi-view and single-view approaches particularly in the presence of partial examples. On single-view benchmarks, extensive evaluation shows that our method is highly competitive or better than other MIML approaches on labeling examples and instances. Moreover, our method can effectively handle datasets with a large number of labels.
Cam-Tu Nguyen, Xiaoliang Wang 0001, Jing Liu 0001, Zhi-Hua Zhou
AAAI2
2014 MIP: Minimizing the idle period of data transmission in data center networks
abstract
In today's data center networks, incast congestion happens when multiple servers send data to one receiver simultaneously. Such congestion results in long idle periods of transmission, which significantly delays the mission complete time (MCT). In this paper, we introduce the MIP, a simple distributed scheme at sender side to Minimize the Idle Periods. To this end, MIP increases the concurrency of data transmission by using proactive fair rate control and cut down idle period further by carefully selected RT O. Extensive simulations show that MIP is able to provide near-optimal MCT performance. In particular, MIP requires no modification to OS kernel or switches in current data center networks.
Chen Deng, Xiaoliang Wang 0001, Sanglu Lu
ICC2
2014 Designing a disaster-resilient network with software defined networking
abstract
With the wide deployment of network facilities and the increasing requirement of network reliability, the disruptive event like natural disaster, power outage or malicious attack has become a non-negligible threat to the current communication network. Such disruptive event can simultaneously destroy all devices in a specific geographical area and affect many network based applications for a long time. Hence, it is essential to build disaster-resilient network for future highly survivable communication services. In this paper, we focus on the integrated approach through the technique of software defined networking to mitigate disaster risks while cut down the investment and management costs. Our design consists of a sub-graph based proactive protection approach for fast rerouting at the network nodes and a splicing approach at the controller for effective post-disaster restoration. Such a systematic design is implemented in OpenFlow framework through the Mininet emulator and Nox controller. Numerical results show that our approach can achieve high reliability, fast recovery and low control overhead.
An Xie, Xiaoliang Wang 0001, Wei Wang 0002, Sanglu Lu
IWQoS2
2014 Delay and Capacity Analysis in MANETs with Correlated Mobility and ${f}$ -Cast Relay
abstract
Many studies have presented the order sense results of information transmission capacity and packet delivery delay in mobile ad hoc networks (MANETs). To achieve the fundamental understanding of MANETs, we focus on deriving the closed-form expressions of the network capacity and end-to-end delay. A MANET with the generalized correlated mobility model is considered in this paper, where the mobility of nodes clustered in one group is confined within a specified area, and multiple groups move uniformly across the network. We also leverage limited packet redundancy to speed up the packet transmission, i.e., each source node is allowed to distribute at most f copies of each packet in its delivery process. Specifically, we first propose an effective multi-hop scheduling-routing scheme under the correlated mobility model, and then develop the closed-form expressions of both per node throughput capacity and expected end-to-end delay. We further explore the tradeoff between throughput capacity and packet delay by using packet redundancy f. The simulation studies validate our theoretical results.
Chen Wang 0009, Xiaoliang Wang 0001, Song Guo 0001, Sanglu Lu
IEEE Trans. Parallel Distributed Syst.3
2013 Recovering erroneous data bits using error estimating code
abstract
Error correction techniques play an important role to guarantee reliable communication in wireless networks. The widely used error-correcting codes (ECCs) such as Hamming code introduce the benefit of error correction without retransmitting the data packet, but they suffer from high redundancy and communication overhead. In the recent years, error estimating code (EEC) was proposed to estimate the bit-error-rate (BER) of a packet efficiently with very low data redundancy. However, the ability of error correction using EEC remains unexplored. In this paper, we argue that EEC can be used to recover erroneous bits from the data packet. To show the capacity of error recovery with EEC, we propose an error correction scheme based on the parity check information provided by the EEC bits. We first introduce a filtering algorithm to rule out the correct data bits and obtain a set of suspicious bits containing most of the errors. Then we apply a polynomial randomized algorithm called Rand_flipping to examine the suspicious bits and flip the most promising erroneous bits aiming to minimize the total numbers of errors in the packet. Theoretical analysis proves that under some constraints the proposed Rand_flipping algorithm can correct most of the erroneous bits with probability higher than 1-1/e. Extensive experiments based on a real WiFi trace are conducted, which shows that the proposed algorithm corrects over 80% erroneous bits of the trace in practice.
Xingshen Wei, Xiaoliang Wang 0001, Sanglu Lu, Xiaoming Fu 0001
ISCC3
2013 Tracing Influential Nodes in a Social Network with Competing Information
Bolei Zhang, Zhuzhong Qian, Xiaoliang Wang 0001, Sanglu Lu
PAKDD (2)3
2013 Capacity and delay of heterogeneous wireless networks with correlated mobility
abstract
Although the capacity of wireless ad hoc networks has been extensively studied under different mobility models and network settings, few work has been done on the effect of heterogeneous mobile nodes and correlated mobility. In this paper, we consider the heterogeneous wireless networks consisting of two types of nodes, called user nodes and master nodes. Specifically, user nodes are combined into groups and each group is equipped with a more powerful master node serving as relay for packet transmissions among groups. By proposing a simple, asymptotically optimal scheduling and routing scheme, we present the maximum per-node throughput and the end-to-end delay in order sense, respectively. We also explore the trade-offs between capacity and delay by adjusting network settings.
Yanzhi Tao, Xiaoliang Wang 0001, Sanglu Lu, Wenbin Jiang 0001
WCNC3
2012 Throughput capacity in mobile ad-hoc networks with correlated mobility and f-cast relay
abstract
The two hop relay algorithms with redundancy are attractive for mobile ad hoc networks (MANET) since they are simple and efficient. In this paper, we extend the analysis of the f-cast two-hop relay algorithm under i.i.d. mobility model to the case of corrected nodes movements, where the source node is allowed to send up to f copies of a packet and the clustered nodes move uniformly across the network. We first provide an effective scheduling-routing algorithm for packet relay inter- and intra-cluster and then explore the scaling laws of throughput capacity of the considered network. This result helps us to study the impact of both the packet redundancy and correlated node movement, and guide us to find the maximum possible throughput capacity through a proper setting of redundancy f.
Chen Wang 0009, Xiaoliang Wang 0001, Sanglu Lu
GLOBECOM3
2012 Assessing physical network vulnerability under random line-segment failure model
abstract
The communication network is now one of the critical infrastructures in our society. However, the current communication networks are facing more and more large-scale region failure threats, such as natural disasters (e.g. earthquake, tornado) and physical attacks (e.g. dragging anchors or EMP attack). Therefore, a deep understanding of network behaviors under region failure is essential for the design and maintenance of future highly survivable networks. In this paper, we focus on the network vulnerability assessment under the geographically correlated region failure(s) caused by a random “line-segment” cut, an important region failure model that can efficiently capture the behaviors of some region failures like earthquake, tornado and anchor cutting. To facilitate such vulnerability assessment, we apply the geometrical probability theory to design a grid partition-based estimation scheme for Disrupted Link Capacity, Pairwise Traffic Reduction and Pairwise Disconnection Probability, three commonly used metrics for statistical vulnerability assessment. A theoretical framework is also established to determine a suitable grid partition such that a specified estimation error requirement is satisfied.
Xiaoliang Wang 0001, Xiaohong Jiang 0001, Achille Pattavina, Sanglu Lu
HPSR1
2012 Exploring social properties in vehicular ad hoc networks
abstract
Vehicular Ad Hoc Networks (VANETs) enable car-to-car communication without the support of network infrastructure, which introduce diverse application possibilities and have drawn much attention from academy and industry in the past years. Unlike other ad hoc networks, nodes in VANETs are restricted to move in streets and have limited communication ranges. Intuitively, vehicle-to-vehicle communication somehow has similarity to human-to-human interaction, which lead to an interesting question of exploring the social properties of VANET nodes. To address the question, we consider encounters of vehicles as their social relationships and model VANETs as social graphs. Based on the social graph model, we use two traces of mobile vehicles from San Francisco and Shanghai to explore their social properties. Our analysis show that several universal laws of social network are hold for VANETs. The social graphs forming by vehicles are scale-free networks with power-law like distribution of node degrees. Small world phenomenon is also observed in our experiments: the nodes in VANETs have high cluster coefficient and there exist short paths between node pairs less than 3 hops on average. The implication of our analytical results is of benefit to develop large scale software system for mobile applications such as VANETs, as well as helps to facilitate inter-device wireless communications in pervasive environment.
Xin Liu 0013, Zhuo Li 0003, Sanglu Lu, Xiaoliang Wang 0001, Daoxu Chen
Internetware5
2012 Constructing N-to-N Shared Optical Queues With Switches and Fiber Delay Lines
abstract
All-optical router has been considered as a natural choice to keep pace with growing fiber link capacity. One main research issue of all-optical router is the design of optical queues with the same flexibility as their electronic counterparts, and some recent works have proved the feasibility of using optical switches and fiber delay lines (SDL) to emulate the electronic queues. In this paper, we focus on the SDL-based construction of N -to- N shared optical queue, a more efficient queue structure in comparison with the dedicated input and output queues. The construction we consider consists of a crossbar switch of size (N + M) × (N + M), where N inputs (outputs) are reserved for external arrivals (departures), and M fiber delay lines are connected from the remaining M outputs back to the remaining M inputs. We first show that by setting the length riof fiber delay line i as ri= 1 + [(1 - 1) mod N], i = 1, . . . M, and scheduling packets properly among these delay lines, such a construction can work as a non- idling first in first out (FIFO) shared queue of size B = Σi=1Mri. We further extend our work to the design of more general shared buffer, where the packets can be stored for an arbitrary time and may depart in a non-FIFO order.
Xiaoliang Wang 0001, Xiaohong Jiang 0001, Achille Pattavina
IEEE Trans. Inf. Theory1
2011 An improved design of optical LIFO buffer with switched delay lines
abstract
The lack of optical buffer is still one of the main problems that hinder the development of all optical networks. One approach to this problem is to emulate the behavior of optical buffers by using optical switches and fiber delay lines (SDL). Current works on this topic have demonstrated the feasibility of constructing SDL-based First In First Out (FIFO) buffer, Priority buffer, etc. The Last In First Out (LIFO) buffer is another important network component for congestion control and QoS guarantee, and parallel and cascade architectures have been peoposed for the efficient design of such optical buffer. The recent work in showed that it is possible to use M fiber delay lines (FDLs) to construct a LIFO buffer of size B = (3/2) · 2M/2- 1 and B = 2(M+1)/2- 1 when M is even and odd, respectively. In this paper, we improve the work in [3] by providing a more efficient construction of SDL-based optical LIFO buffer. We first show that with a single stage feedback structure consisted of one (M + 1) × (M + 1) crossbar switch and M FDLs connecting M outputs of the crossbar back to M its inputs, we are able to construct a LIFO buffer of size B = 2 · 2M/2- 2 and B = (3/2) · 2(M+1)/2- 2 when M is even and odd, respectively. This is achieved through adopting a properly delay length setting for each FDL and a careful packets scheduling among FDLs, as well as exploiting the nice function of simultaneous packet reading and witting a FDL can support. We further show that if we adopt a cascade of smaller switches rather than a single (M+1)×(M+1) big switch, the new LIFO design can be implemented with much lower complexity in terms of the total number of basic 2 × 2 switch elements.
Xiaoliang Wang 0001, Xiaohong Jiang 0001, Achille Pattavina
HPSR1
2011 Assessing network vulnerability under probabilistic region failure model
abstract
The mission critical network infrastructures are facing potential large region threats, both intentional (like EMP attack, bomb explosion) and natural (like earthquake, flooding). The available research on region failure related vulnerability studies generally adopt a kind of simple “deterministic” region failure models, which can not capture some important features of real region failure scenarios, where a network component in the region only fails with certain probability, and more importantly, such failure probability tends to vary with both its dimension and its distance to failure center. In this paper, we provide a more general “probabilistic” region failure model to capture the key features of a region failure and apply it for the network vulnerability assessment. To facilitate such assessment, we adopt a grid partition-based scheme to estimate various statistical network metrics under a random region failure. A theoretical framework is also established to determine a suitable grid partition such that a specified estimation error requirement is satisfied. The grid partition technique is also useful for identifying the vulnerable zones of a network, which can guide network designers to initiate proper network protection against such failures. The work in this paper helps us more deeply understand the network vulnerability behavior under region failure and facilitates the design and maintenance of future highly survivable mission critical networks.
Xiaoliang Wang 0001, Xiaohong Jiang 0001, Achille Pattavina
HPSR1
2011 Efficient Designs of Optical LIFO Buffer with Switches and Fiber Delay Lines
abstract
The lack of optical buffer is still one of the main problems that hinder the development of all optical networks. Current works on this topic mainly focus on the emulation of optical buffers based on a combination of fiber delay line (FDL) and switch. These works have demonstrated the feasibility of FDL-based emulation for many kinds of optical buffers, like the First In First Out (FIFO) buffer, Priority buffer, etc. The Last In First Out (LIFO) buffer is another basic network component for congestion control and QoS guarantee. Recently, Huang et al. introduced a recursive construction for the LIFO buffer, which requires no less than 9 log2B FDLs to build such a buffer of size B. In this paper, we first show that by a proper FDL grouping and a suitable FDL length assignment for each FDL-group, we need approximately 3 log2B FDLs to emulate a LIFO buffer of size B. We then demonstrate that if a careful packet scheduling among FDL-groups is adopted, this number of FDLs can be further reduced to 2 log2B.
Xiaoliang Wang 0001, Xiaohong Jiang 0001, Achille Pattavina
IEEE Trans. Commun.1
2009 A construction of 1-to-2 shared optical buffer queue with switched delay lines
abstract
Optical buffering is fundamental to contention resolution in optical networks. The current works on this line mainly focus on the emulation of dedicated input/output buffer queue by using switched fiber delay lines (SDL). It is notable that the shared buffer queue, where a common buffer pool is shared by all the input/output ports of a switch, has the potential to significantly reduce the overall buffer capacity requirement. As far as we know, however, no related work is available yet on the exact emulation of a shared optical buffer queue with SDLs. In this paper, we focus on the design of first in first out (FIFO) shared optical buffer queue based on the optical feedback SDL construction. The construction considered consists of an (M + 2) x (M + 2) switch fabric and M fiber delay lines FDL1, . . ., FDLM, where FDLiconnects the ithoutput of the switch fabric with its ithinput. We show that by setting the length of FDL1as min(M + 1 - i, i), i = 1, . . . , M, such a construction can actually work as an 1-to-2 shared buffer queue. We then extend this emulation to the more general N-to-2 case.
Xiaoliang Wang 0001, Xiaohong Jiang 0001, Achille Pattavina, Susumu Horiguchi
IEEE Trans. Commun.1
2008 Maintaining Packet Order in Reservation-Based Shared-Memory Optical Packet Switch
abstract
Shared-memory optical packet (SMOP) switch architecture is very promising for significantly reducing the amount of required optical memory, which is typically constructed from fiber delay lines (FDLs). The current reservation-based scheduling algorithms for SMOP switches can effectively utilize the FDLs and achieve a low packet loss rate by simply reserving the departure time for each arrival packet. It is notable, however, that such a simple scheduling scheme may introduce a significant packets out of order problem. In this paper, we first identify the two main sources of packets out of order in the current reservation-based SMOP switches. We then show that by introducing a "last-timestamp " variable and modifying the corresponding FDLs arrangement as well as the scheduling process in the current reservation-based SMOP switches, it is possible to keep packets in-sequence while still maintaining a similar delay and packet loss performance as the previous design.
Xiaoliang Wang 0001, Xiaohong Jiang 0001, Susumu Horiguchi
AINA1
2007 CBX-1 Switch: An Effective Load Balanced Switch
abstract
Load balance (LB) switch architecture is attractive for building high speed switches since it does not require a centralized scheduler and still can guarantee 100% throughput under any admissible traffic. It is notable, however, that due to the multi-path property of LB switches, the packet out-of-sequence problem may happen in such switches. Several schemes have been proposed to tackle the above problem, but these schemes either require infinite central buffer or introduce a high average packet delay. In this paper, we propose a new LB switch architecture - central buffer one-packet-crosspoint (CBX-1). The key idea ofCBX-1 is to introduce a VIOQ (virtual input output queue) with capacity one (i.e., it can store only one packet) after the first stage of a LB switch to emulate a CIXB-1 switch (combined input-one-packet-crosspoint buffered switch). We show through both analysis and simulation that although our architecture requires only finite central buffer to tackle the packet out-of-sequence problem, it still guarantees 100% throughput and achieves a good delay performance.
Xiaoliang Wang 0001, Xiaohong Jiang 0001, Susumu Horiguchi
PDCAT1