Dan Li 0001

dblp:48/4185-1 · DBLP profile ↗
← Back
130ranked-venue papers
22as first author
59since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 92 · 16 first-author · 46 since 2021Systems, architecture and hardware · 18 · 4 first-author · 3 since 2021Security and privacy · 7 · 7 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 OptiFlow: Towards LLM-Driven Optimization of Collective Communication Algorithms
Ziyue Yang 0002, Kaihui Gao, Shuai Wang 0028, Li Chen 0008, Zhixiong Niu, Ran Shu 0001, Wenxue Cheng, Peng Cheng 0005, Yongqiang Xiong, Dan Li 0001
APNet11
2026 CATS: Predictive-Feedback Adaptive Load Balancing for Computing-Aware Traffic Steering
Yuxiang Shang, Tao Sun 0010, Dan Li 0001, Zhenping Hu, Lu Lu 0016, Chengjiang Wen, Yantao Han, Li Chen 0008, Huijuan Yao, Peng Liu 0047
ICC3
2026 Accurate and Stable AS Relationship Inference via Trusted Seeds and Semi-Supervised Learning
Siyuan Teng, Lancheng Qin, Li Chen 0008, Dan Li 0001
INFOCOM4
2026 A Large-Scale IPv6-Based Measurement of the Starlink Network
abstract
Low Earth Orbit (LEO) satellite networks have attracted considerable attention for their ability to deliver global, low-latency broadband Internet services. In this paper, we present a large-scale measurement study of the Starlink network, the largest LEO satellite constellation to date. We first propose an efficient method for discovering active Starlink user routers, identifying approximately 5.98 million IPv6 addresses across 208 regions in 165 countries. Compared to general-purpose IPv6 target generation algorithms, our router-centric approach achieves near-complete coverage and, to the best of our knowledge, yields the most comprehensive known set of active IPv6 addresses for Starlink user routers. Based on the discovered user routers, we further propose an efficient method for mapping the Starlink backbone network and uncover a topology consisting of 49 Points of Presence (PoPs) interconnected by 98 links. We conduct a detailed statistical analysis of active Starlink user routers and PoPs, and further characterize the IPv6 address assignment strategy adopted by the Starlink network. Finally, we analyze the latency of Starlink user routers, propose a method to distinguish different types of users within the same region using outside-in measurement, and identify the ongoing V2 Mini satellite deployment as a potential driver of the performance improvements. The dataset of the Starlink backbone network is publicly available at https://ki3.org.cn/#/starlink-network.
Bingsen Wang, Shuai Wang 0028, Li Chen 0008, Jinwei Zhao, Dan Li 0001, Yong Jiang 0001
INFOCOM6
2026 HeraClass: Towards Open-World Network Flow Classification via Traffic-Language Mapping
Ni Jin, Libin Liu 0001, Yukai Miao, Li Chen 0008, Dan Li 0001, Xizheng Wang, Xiuting Xu, Baojiang Cui
IWQoS5
2026 OSAVRoute: Advancing Outbound Source Address Validation Deployment Detection with Non-Cooperative Measurement
Shuai Wang 0028, Li Chen 0008, Dan Li 0001, Lancheng Qin
NDSS4
2026 BayWatch: Practical Internet-Scale Topology Monitoring with Dynamic Bayesian Estimation
Zhongxu Guan, Shuai Wang 0028, Li Chen 0008, Zhaoteng Yan, Jiaye Lin, Dan Li 0001, Yong Jiang 0001, Yingxin Wang
NSDI6
2026 CCEval: Accurately and Confidently Evaluating Performance Metrics of Congestion Control Algorithms for Datacenter Networks
Tianfeng Liu, Kaihui Gao, Li Chen 0008, Dan Li 0001, Jin Guang, Vincent Liu 0001, Yiwei Zhang 0016, Ni Jin
NSDI4
2026 Supercharging Packet-level Network Simulation of Large Model Training via Memoization and Fast-Forwarding
Kaihui Gao, Li Chen 0008, Dan Li 0001, Yiwei Zhang 0016, Fei Gui, Yitao Xing, Wenjia Wei, Bingyang Liu
NSDI4
2026 PReCCL: Performant and Resilient Collective Communication via Integrated Inband Telemetry and Workload Reallocation
abstract
Modern collective communication libraries (CCLs) execute a collective communication task (CCT) by decomposing it into multiple sub-tasks, each mapped to a specific Virtual Topology (VT), which is an ordered graph of GPUs (e.g., a ring or a tree), to maximize parallelism and link utilization. As AI training scales to larger clusters, network anomalies (congestion and failures) are unavoidable, and a single straggling VT can delay the entire CCT. Existing solutions either rely on low-level transport-layer solutions which lacks a cross-sub-task perspective, or static CCL scheduling, failing to adapt to the dynamic and heterogeneous networks.
Kaihui Gao, Li Chen 0008, Fei Gui, Dan Li 0001, Jiamin Cao
SIGCOMM7
2026 Networked Agent Memory and Causality Representation: Experiences towards Interpretable Cloud-Scale Root-Causing
Yanyu Ren, Xianshang Lin, Chenxu Wang 0007, Li Chen 0008, Shuai Wang 0028, Kaihui Gao, Dan Li 0001, Chen Tian 0001, Yunguang Li, Ennan Zhai
SIGCOMM7
2026 Open the Floodgates in a Digital Twin: Experiences of Building Spillway for 100M+-User Signaling Storms in Cellular Core Network
abstract
Signaling storms threaten cellular core networks when synchronized reconnection attempts from massive numbers of devices trigger cascading, metastable overloads. Existing defenses rely on manual, static configurations of local overload controls, which ignore serial dependencies among heterogeneous network elements. We present Spillway, a digital-twin-driven system that automates global signaling-flood mitigation. Spillway introduces a hierarchical defense architecture that enforces altruistic throttling, allowing upstream nodes to shed load before downstream bottlenecks collapse. To evaluate candidate configurations, Spillway uses CN-DES, a domain-specific discrete-event simulator with a vectorized kernel. By aggregating users that share protocol states, CN-DES decouples simulation cost from user count and simulates regional-scale storms involving tens of millions of users in minutes, achieving a 60× speedup over traditional simulation while preserving fidelity. Spillway then uses heteroscedastic evolutionary Bayesian optimization to search a large, non-convex parameter space. We report on a five-year deployment in the world's largest 5G Standalone network. During real incidents, including application anomalies and RAN failures, networks using Spillway-optimized configurations experienced substantially fewer user fallbacks than predicted under legacy configurations; post-incident analysis confirms that pre-deployed parameters kept all network elements within safe operating bounds.
Hongtao Xie 0006, Jianmin Liu, Li Chen 0008, Dan Li 0001, Mineng Fu, Xi Chen 0026
SIGCOMM6
2026 I2BGP: A Privacy-Preserving Intra-AS State-Assisted Inter-AS Routing Scheme
abstract
BGP is the most widely employed inter-AS routing protocol, connecting millions of ASes worldwide. While it is possible to select the egress for outgoing flows based on administrators’ configurations, such schemes are localized due to the privacy of the intra-AS network state. TheAS_Pathfield of BGP records all crossed ASes, which can be used to prevent routing loops and select paths,i.e., selecting the minimum AS-hop path among available paths. Although this scheme is simple, effective, and offers a certain degree of global perspective, selecting paths at AS granularity ignores the transmission performance within each intra-AS, which may result in selecting non-optimal routing paths. To enable the use of private intra-AS data for inter-AS routing, we proposed a privacy-preserving intra-AS state-assisted inter-AS routing scheme, which can select optimal inter-AS paths without disclosing specific intra-AS state data. Specifically, we added an additional BGP header field to carry path performance features and designed a three-step data masking mechanism to protect intra-AS state data, enabling the selection of inter-AS paths with intra-AS state awareness. I2BGP has been deployed in the Greater Bay Area Future Network and a large-scale network simulator based on real network topologies. The results show that I2BGP outperforms BGP in terms of specified forwarding hops, delay, and bandwidth metrics.
Peizhuang Cong, Yuchao Zhang 0004, Jun Wang 0178, Wendong Wang 0003, Tong Yang 0003, Dan Li 0001, Ke Xu 0002
IEEE Trans. Netw.6
2026 Example Generalizing Network Configuration Synthesizer via Graph-Informed Large Language Models
Jianmin Liu, Li Chen 0008, Dan Li 0001, Yukai Miao, Liyu Ma
IEEE Trans. Netw.3
2025 A Theoretical Framework for Quantitative Evaluation of Padding Defenses against Website Fingerprinting
Dan Li 0001
APNet3
2025 Towards Automatic Network Diagram Comprehension
abstract
Network Diagram Comprehension (NDC) is a vital task for networking professionals, offering essential insights into network topology and configurations. However, NDC remains a labor-intensive process heavily reliant on human expertise, with existing tools falling short in addressing this challenge. It is critical to develop an Automatic NDC (ANDC) system that ensures high faithfulness and completeness in information extraction while supporting practical, end-to-end NDC applications. Moreover, a comprehensive dataset and benchmark are necessary to systematically evaluate and drive the progress of ANDC.In this work, we introduce Layered Extractor of Network Diagrams (LEND), the first ANDC system designed to comprehensively and faithfully extract and utilize information from network diagrams. LEND employs a three-stage pipeline: (1) a layer extractor to decompose diagrams and identify key elements with a denoising cascade, (2) an inter-layer combiner to reconstruct entity relations with positional and domain knowledge, and (3) a task-specific interpreter for networking applications.To support this effort, we develop two extensive NDC datasets comprising over 4,000 network diagrams and icons from diverse sources, along with the first benchmark to evaluate ANDC systems across three distinct metrics. Empirical experiments demonstrate that LEND outperforms existing methods by achieving at 1.21– 5.10× better faithfulness and completeness, and improves its capability as a NetOps engineer by 30.5% on the Cisco Certified Network Associate (CCNA) exam.
Yanyu Ren, Yukai Miao, Li Chen 0008, Dan Li 0001, Xizheng Wang, Yu Bai 0021
ICNP4
2025 Measuring the Time Source Vulnerabilities in the NTP Ecosystem
abstract
Precise timekeeping is crucial for the dependable functioning and security of multiple Internet infrastructures, such as TLS certificates. Although the Network Time Protocol (NTP) is widely used for time synchronization across devices, it has several security vulnerabilities. Network Time Security (NTS) offers server authentication and integrity verification to protect against man-in-the-middle attacks. However, NTS does not address issues related to erroneous time sources.
Zhentian Huang, Shuai Wang 0028, Li Chen 0008, Dan Li 0001
IMC4
2025 SAIP: Accurate Detection of Anycast Servers with the Rise of Regional Anycast
Shuai Wang 0028, Li Chen 0008, Dan Li 0001
INFOCOM4
2025 Transcending Cost-Quality Tradeoff in Agent Serving via Session-Awareness
abstract
Large Language Model (LLM) agents are capable of task execution across various domains by autonomously interacting with environments and refining LLM responses based on feedback. However, existing model serving systems are not optimized for the unique demands of serving agents. Compared to classic model serving, agent serving has different characteristics: predictable request pattern, increasing quality requirement, and unique prompt formatting. We identify a key problem for agent serving: LLM serving systems lack session-awareness. They neither perform effective KV cache management nor precisely select the cheapest yet competent model in each round. This leads to a cost-quality tradeoff, and we identify an opportunity to surpass it in an agent serving system. To this end, we introduce AgServe for AGile AGent SERVing. AgServe features a session-aware server that boosts KV cache reuse via Estimated-Time-of-Arrival-based eviction and in-place positional embedding calibration, a quality-aware client that performs session-aware model cascading through real-time quality assessment, and a dynamic resource scheduler that maximizes GPU utilization. With AgServe, we allow agents to select and upgrade models during the session lifetime, and to achieve similar quality at much lower costs, effectively transcending the tradeoff. Extensive experiments on real testbeds demonstrate that AgServe (1) achieves comparable response quality to GPT-4o at a 16.5\% cost. (2) delivers 1.8$\times$ improvement in quality relative to the tradeoff curve.
Yanyu Ren, Li Chen 0008, Dan Li 0001, Xizheng Wang, Yukai Miao, Yu Bai 0021
NeurIPS3
2025 Accelerating Design Space Exploration for LLM Training Systems with Multi-experiment Parallel Simulation
Fei Gui, Kaihui Gao, Li Chen 0008, Dan Li 0001, Vincent Liu 0001, Hongbing Yang, Dian Xiong
NSDI4
2025 CEGS: Configuration Example Generalizing Synthesizer
Jianmin Liu, Li Chen 0008, Dan Li 0001, Yukai Miao
NSDI3
2025 Resolving Packets from Counters: Enabling Multi-scale Network Traffic Super Resolution via Composable Large Traffic Model
Xizheng Wang, Libin Liu 0001, Li Chen 0008, Dan Li 0001, Yukai Miao, Yu Bai 0021
NSDI4
2025 SimAI: Unifying Architecture Design and Performance Tuning for Large-Scale Large Language Model Training with Scalability and Precision
Xizheng Wang, Qingxu Li, Yichi Xu, Dan Li 0001, Li Chen 0008, Heyang Zhou, Linkang Zheng, Yikai Zhu, Yang Liu 0245, Kun Qian 0021, Kunling He, Ennan Zhai, Dennis Cai, Binzhang Fu
NSDI5
2025 Discovering Millions of New Nodes and Links in the Internet by Challenging the Uniformity Assumption in Multipath Detection
abstract
Multipath Detection Algorithms (MDAs) are proposed to discover Internet topology in the presence of load balancing (LB). Existing methods assume uniformity in the load-balancing responses (LBR), i.e., responses from the successors of a LB router. However, we reveal that only 20% of the cases exhibit uniformity in the Internet. This finding significantly challenges the completeness of the Internet topology discovered using current MDAs. In this paper, we propose a novel system BayMuDA, that can estimate LBR distributions and calculate the minimum number of probes needed to statistically discover all nodes and links within a given hop. The validation on controlled topologies shows that BayMuDA discovers at least 85%/73% of nodes/links in ~90% of the cases. Our Internet-wide measurement results indicate that BayMuDA can discover millions of Internet nodes and links obscured by the state-of-the-art MDA algorithm, D-Miner, due to uneven responses.
Zhongxu Guan, Shuai Wang 0028, Li Chen 0008, Zhaoteng Yan, Jiaye Lin, Dan Li 0001, Yong Jiang 0001, Yingxin Wang
SIGCOMM6
2025 From ATOP to ZCube: Automated Topology Optimization Pipeline and A Highly Cost-Effective Network Topology for Large Model Training
abstract
The development of large language models (LLMs) poses new challenges in data center network topology design. To assist in exploring topology design, we propose ATOP, an Automated Topology Optimization Pipeline, which models network topology as a set of hyperparameters, enabling the discovery of potential topologies. With various optimization algorithms and customizable optimization objectives, ATOP achieves automated topology optimization on a scale of tens of thousands of GPUs. We apply ATOP on network topologies for 256, 1024, 4096, and 16384 GPUs, optimizing performance under LLMs training traffic patterns, collective communication performance, fault tolerance, and network cost. We also evaluate ATOP in different scenarios: building, optimizing, and expanding a data center. From ATOP's results, we discover a new topology — ZCube, which reaches the highest cost-effectiveness across various GPU scales. Simulation results show that ZCube, compared to the previous state-of-the-art topologies, including Rail-optimized Fat-tree (ROFT), Rail-only, and HPN, improves end-to-end LLM training speed by 3% to 7% and reduces network hardware costs by 26% to 46%. We also construct ZCube on a real-world testbed. Results show that ZCube reduces hardware costs by 25% compared to Rail-Optimized Topology while maintaining the same all-reduce and all-to-all performance.
Dan Li 0001, Li Chen 0008, Dian Xiong, Kaihui Gao, Yiwei Zhang 0016, Menglei Zhang, Bochun Zhang, Zhuo Jiang, Jianxi Ye, Haibin Lin
SIGCOMM2
2025 The Digital Cybersecurity Expert: How Far Have We Come?
abstract
The increasing deployment of large language models (LLMs) in the cybersecurity domain underscores the need for effective model selection and evaluation. However, traditional evaluation methods often overlook specific cybersecurity knowledge gaps that contribute to performance limitations. To address this, we develop CSEBenchmark, a fine-grained cyber-security evaluation framework based on 345 knowledge points expected of cybersecurity experts. Drawing from cognitive science, these points are categorized into factual, conceptual, and procedural types, enabling the design of 11,050 tailored multiple-choice questions. We evaluate 12 popular LLMs on CSEBenchmark and find that even the best-performing model achieves only 85.42% overall accuracy, with particular knowledge gaps in the use of specialized tools and uncommon commands. Different LLMs have unique knowledge gaps. Even large models from the same family may perform poorly on knowledge points where smaller models excel. By identifying and addressing specific knowledge gaps in each LLM, we achieve up to an 84% improvement in correcting previously incorrect predictions across three existing benchmarks for two cybersecurity tasks. Furthermore, our assessment of each LLM's knowledge alignment with specific cybersecurity roles reveals that different models align better with different roles, such as GPT-4o for the Google Senior Intelligence Analyst and Deepseek-V3 for the Amazon Privacy Engineer. These findings underscore the importance of aligning LLM selection with the specific knowledge requirements of different cybersecurity roles for optimal performance.
Dawei Wang 0021, Geng Zhou, Yu Bai 0021, Li Chen 0008, Ting Qin, Dan Li 0001
SP8
2025 Your Shield is My Sword: A Persistent Denial-of-Service Attack via the Reuse of Unvalidated Caches in DNSSEC Validation
Shuai Wang 0028, Li Chen 0008, Dan Li 0001
USENIX Security Symposium4
2024 ProphetFuzz: Fully Automated Prediction and Fuzzing of High-Risk Option Combinations with Only Documentation via Large Language Model
abstract
Vulnerabilities related to option combinations pose a significant challenge in software security testing due to their vast search space. Previous research primarily addressed this challenge through mutation or filtering techniques, which inefficiently treated all option combinations as having equal potential for vulnerabilities, thus wasting considerable time on non-vulnerable targets and resulting in low testing efficiency. In this paper, we utilize carefully designed prompt engineering to drive the large language model (LLM) to predict high-risk option combinations (i.e., more likely to contain vulnerabilities) and perform fuzz testing automatically without human intervention. We developed a tool called ProphetFuzz and evaluated it on a dataset comprising 52 programs collected from three related studies. The entire experiment consumed 10.44 CPU years. ProphetFuzz successfully predicted 1748 high-risk option combinations at an average cost of only \8.69 per program. Results show that after 72 hours of fuzzing, ProphetFuzz discovered 364 unique vulnerabilities associated with 12.30% of the predicted high-risk option combinations, which was 32.85% higher than that found by state-of-the-art in the same timeframe. Additionally, using ProphetFuzz, we conducted persistent fuzzing on the latest versions of these programs, uncovering 140 vulnerabilities, with 93 confirmed by developers and 21 awarded CVE numbers.
Dawei Wang 0021, Geng Zhou, Li Chen 0008, Dan Li 0001, Yukai Miao
CCS4
2024 Robust or Risky: Measurement and Analysis of Domain Resolution Dependency
abstract
DNS relies on domain delegation for good scalability, where domains delegate their resolution service to authoritative nameservers. However, such delegations lead to complex inter-dependencies between DNS zones. While a complex dependency might improve the robustness of domain resolution, it could also introduce security risks unexpectedly. In this work, we perform a large-scale measurement on nearly 217M domains to analyze their resolution dependencies at both zone level and infrastructure level. According to our analysis, domains under country-code TLDs and new generic TLDs generally present more complicated dependency relationships. For robustness consideration, popular domains prefer to configure more complex dependencies. However, the centralization of nameserver hosting and the silent outsourcing of DNS providers could lead to severe false redundancy at infrastructure level. Worse, considerable domain configurations in the wild are "not robust but risky": a more complex dependency may also indicate more vulnerabilities, e.g., domains with a 2× higher dependency complexity have a 2.87× larger proportion suffering from the hijacking risk brought by lame delegation.
Shuai Wang 0028, Dan Li 0001
INFOCOM3
2024 Understanding Route Origin Validation (ROV) Deployment in the Real World and Why MANRS Action 1 Is Not Followed
Lancheng Qin, Li Chen 0008, Dan Li 0001, Honglin Ye
NDSS3
2024 dRR: A Decentralized, Scalable, and Auditable Architecture for RPKI Repository
Dan Li 0001, Li Chen 0008, Qi Li 0002, Sitong Ling
NDSS2
2024 RedTE: Mitigating Subsecond Traffic Bursts with Real-time and Distributed Traffic Engineering
abstract
Internet traffic bursts usually happen within a second, thus conventional burst mitigation methods ignore the potential of Traffic Engineering (TE). However, our experiments indicate that a TE system, with a sub-second control loop latency, can effectively alleviate burst-induced congestion. TE-based methods can leverage network-wide tunnel-level information to make globally informed decisions (e.g., balancing traffic bursts among multiple paths). Our insight in reducing control loop latency is to let each router make local TE decisions, but this introduces the key challenge of minimizing performance loss compared to centralized TE systems.
Fei Gui, Dan Li 0001, Li Chen 0008, Kaihui Gao, Congcong Min, Yi Wang 0004
SIGCOMM3
2024 Deep Distributional Reinforcement Learning-Based Adaptive Routing With Guaranteed Delay Bounds
abstract
Real-time applications that require timely data delivery over wireless multi-hop networks within specified deadlines are growing increasingly. Effective routing protocols that can guarantee real-time QoS are crucial, yet challenging, due to the unpredictable variations in end-to-end delay caused by unreliable wireless channels. In such conditions, the upper bound on the end-to-end delay, i.e., worst-case end-to-end delay, should be guaranteed within the deadline. However, existing routing protocols with guaranteed delay bounds cannot strictly guarantee real-time QoS because they assume that the worst-case end-to-end delay is known and ignore the impact of routing policies on the worst-case end-to-end delay determination. In this paper, we relax this assumption and propose DDRL-ARGB, an Adaptive Routing with Guaranteed delay Bounds using Deep Distributional Reinforcement Learning (DDRL). DDRL-ARGB adopts DDRL to jointly determine the worst-case end-to-end delay and learn routing policies. To accurately determine worst-case end-to-end delay, DDRL-ARGB employs a quantile regression deep Q-network to learn the end-to-end delay cumulative distribution. To guarantee real-time QoS, DDRL-ARGB optimizes routing decisions under the constraint of worst-case end-to-end delay within the deadline. To improve traffic congestion, DDRL-ARGB considers the network congestion status when making routing decisions. Extensive results show that DDRL-ARGB can accurately calculate worst-case end-to-end delay, and can strictly guarantee real-time QoS under a small tolerant violation probability against two state-of-the-art routing protocols.
Jianmin Liu, Dan Li 0001, Yongjun Xu 0001
IEEE/ACM Trans. Netw.2
2024 Decentralized and Incentivized Federated Learning: A Blockchain-Enabled Framework Utilising Compressed Soft-Labels and Peer Consistency
abstract
Federated Learning (FL) has emerged as a powerful paradigm in Artificial Intelligence, facilitating the parallel training of Artificial Neural Networks on edge devices while safeguarding data privacy. Nonetheless, to encourage widespread adoption, Federated Learning Frameworks (FLFs) must tackle (i) the power imbalance between a central authority and its participants, and (ii) the challenge of equitably measuring and incentivizing contributions. Existing approaches to decentralize and incentivize FL processes are hindered by (i) computational overhead and (ii) uncertainty in contribution assessment [1]), limiting FL's scalability beyond use cases where trust between participants and the server is established. This work introduces a cutting-edge, blockchain-enabled federated learning framework that incorporates Federated Knowledge Distillation (FD) with compressed 1-bit soft-labels, aggregated through a smart contract. Furthermore, we present the Peer Truth Serum for Federated Distillation (PTSFD), which cultivates an incentive-compatible ecosystem by rewarding honest participation based on an implicit yet effective comparison of worker contributions. The primary innovation stems from its lightweight architecture that simultaneously promotes decentralization and incentivization, addressing critical challenges in contemporary FL approaches.
Leon Witt, Usama Zafar, KuoYeh Shen, Felix Sattler, Dan Li 0001, Wojciech Samek
IEEE Trans. Serv. Comput.5
2023 sRDMA: A General and Low-Overhead Scheduler for RDMA
abstract
Remote Direct Memory Access (RDMA) has been widely deployed in data centers to improve application performance. However, the characteristic of RDMA to deliver messages in order cannot meet the emerging requirements of applications for scheduling messages within an RDMA connection, making RDMA unable to be fully utilized. Some works try to schedule the data to be transferred in specific applications before delivering to RDMA, or distribute messages to different connections. However, these approaches tightly couple scheduling logic with application logic and may result in high scheduling overhead.
Xizheng Wang, Shuai Wang 0028, Dan Li 0001
APNet3
2023 Impact of International Submarine Cable on Internet Routing
Honglin Ye, Shuai Wang 0028, Dan Li 0001
INFOCOM3
2023 BGL: GPU-Efficient GNN Training by Optimizing Graph Data I/O and Preprocessing
Tianfeng Liu, Yangrui Chen, Dan Li 0001, Chuan Wu 0001, Yibo Zhu 0001, Yanghua Peng, Hongzheng Chen, Chuanxiong Guo
NSDI3
2023 Demo: NetVision: Efficient Visualization Front-End for Packet-level Discrete-Event Network Simulation
abstract
Visualization of network simulation is an essential tool for network practitioners. However, the front-end of existing network simulators often fails to deliver satisfactory performance when dealing with modern network scales and interface speed. In this paper, we propose NetVision, an efficient visualization front-end of network simulation based on the Unity engine, which is commonly used for video game and virtual reality development. NetVision offers flow-level visualization of network behavior and performances. Then, through parallel optimization, NetVision supports real-time visualization for large-scale high-speed networks.
Kaihui Gao, Li Chen 0008, Dan Li 0001, Vincent Liu 0001, Xizheng Wang, Lu Lu 0016
SIGCOMM3
2023 DONS: Fast and Affordable Discrete Event Network Simulation with Automatic Parallelization
abstract
Discrete Event Simulation (DES) is an essential tool for network practitioners. Unfortunately, existing DES simulators cannot achieve satisfactory performance at the scale of modern networks. Recent work has attempted to address these challenges by reducing the traffic processed via novel approximation techniques; however, we argue in this paper that much of the slowdown of existing DES simulators is due to their underlying software architecture.
Kaihui Gao, Li Chen 0008, Dan Li 0001, Vincent Liu 0001, Xizheng Wang, Lu Lu 0016
SIGCOMM3
2023 Poster: Q-Scanner: A Fast Scanning Tool for Large-Scale SSL/TLS Configurations Measurement
abstract
Secure Sockets Layer (SSL) and Transport Layer Security (TLS) protocols are used to encrypt data, protect privacy, and authenticate. However, the security of SSL/TLS itself depends on its configurations. While some scanning tools are used to measure SSL/TLS configurations, their performance is far from meeting the requirement of large-scale measurements. In this paper, we propose a fast SSL/TLS configuration scanning tool, Q-Scanner, which can generate a lightweight scanning solution based on the characteristics of the configurations to be scanned. The experiment shows Q-Scanner achieves a speedup of over 30,000 times compared to SSL Pulse without loss of accuracy.
Shuai Wang 0028, Dan Li 0001
SIGCOMM3
2023 Light: A Compatible, high-performance and scalable user-level network stack
Dan Li 0001, Huiyou Jiang, Du Lin, Jinkun Geng, K. K. Ramakrishnan, Kai Zheng 0003
Comput. Networks2
2023 DIT and Beyond: Interdomain Routing With Intradomain Awareness for IIoT
abstract
Along with the ever-increasing amount of data generated from industrial devices, the cross domain [also known as autonomous systems (ASs)] data transmission problem has attracted more and more attention in the Industrial Internet of Things (IIoT). As mature and widely used interdomain routing protocols, border gateway protocol-based solutions often take the number of domains (i.e., AS hops) of each path as a criterion to make routing decisions, which is simple and effective. However, such protocols can only meet the reachability requirements while ignoring the performance requirements. That is, the path with the minimum AS hops will be selected to carry flows, even if the actual performance of this path does not meet the transmission requirements due to the unawareness of intradomain information on that path. But it is not impractical to directly access intradomain information for making better routing decisions given data privacy concerns. In this article, we propose M-DIT, which can make interdomain routing decisions with the assistance of desensitized intradomain information for multiple-requirement transmissions. To do so, we design a homomorphic encrypted-based private number comparison scheme to export intradomain information securely and, thus, assist in routing decisions. The results of some experiments based on five real topologies (ATMnet,Claranet,Compuserve,NSFnet, andPeer1) with thousands of interdomain flows demonstrate that M-DIT reduced flow completion time by about 60% or selected high bandwidth paths flexibly for interdomain routing for IIoT scenarios.
Peizhuang Cong, Yuchao Zhang 0004, Wendong Wang 0003, Xiangyang Gong, Tong Yang 0003, Dan Li 0001, Ke Xu 0002
IEEE Internet Things J.7
2023 Decentral and Incentivized Federated Learning Frameworks: A Systematic Literature Review
abstract
The advent of federated learning (FL) has sparked a new paradigm of parallel and confidential decentralized machine learning (ML) with the potential of utilizing the computational power of a vast number of Internet of Things (IoT), mobile, and edge devices without data leaving the respective device, thus ensuring privacy by design. Yet, simple FL frameworks (FLFs) naively assume an honest central server and altruistic client participation. In order to scale this new paradigm beyond small groups of already entrusted entities toward mass adoption, FLFs must be: 1) truly decentralized and 2) incentivized to participants. This systematic literature review is the first to analyze FLFs that holistically apply both, the blockchain technology to decentralize the process and reward mechanisms to incentivize participation. 422 publications were retrieved by querying 12 major scientific databases. After a systematic filtering process, 40 articles remained for an in-depth examination following our five research questions. To ensure the correctness of our findings, we verified the examination results with the respective authors. Although having the potential to direct the future of distributed and secure artificial intelligence, none of the analyzed FLFs is production ready. The approaches vary heavily in terms of use cases, system design, solved issues, and thoroughness. We provide a systematic approach to classify and quantify differences between FLFs, expose limitations of current works and derive future directions for research in this novel domain.
Leon Witt, Mathis Heyer, Kentaroh Toyoda, Wojciech Samek, Dan Li 0001
IEEE Internet Things J.5
2023 Buffer-Based High-Coverage and Low-Overhead Request Event Monitoring in the Cloud
abstract
Request latency directly affects the performance of modern cloud applications. Due to various causes in hosts and networks, requests can suffer from request latency anomalies (RLAs), which may violate the Service-Level Agreement. However, existing performance monitoring tools have incomplete coverage and inconsistent semantics for monitoring requests and cannot accurately diagnose RLAs. This paper presentsBufScope, a high-coverage and low-overhead request event monitoring system, which monitorsbuffersto capture most RLA-related abnormal events with consistent request-level semantics in the end-to-end datapath of request. First,BufScopemodels the datapath of request as a buffer chain and defines events based on three properties of buffers, so as toend-to-end monitorthe root causes of RLA. Then, to achieveconsistent semanticsfor captured events,BufScopedesigns a request-level semantics injection mechanism to make events captured in networks have the victim requests’ ID. Finally,BufScopeoffloads the semantics operations and event collection in software to SmartNICs forlow CPU overhead. We have implementedBufScopeon commodity SmartNICs and programmable switches. Evaluation results show thatBufScopecan diagnose 98% RLAs with < 0.08% network bandwidth overhead and 0.6% application throughput decline.
Kaihui Gao, Chen Sun 0005, Shuai Wang 0028, Dan Li 0001, Yu Zhou 0008, Hongqiang Harry Liu, Lingjun Zhu, Ming Zhang 0005, Lu Lu 0016
IEEE/ACM Trans. Netw.4
2023 Dependable Virtualized Fabric on Programmable Data Plane
abstract
In modern multi-tenant data centers, each tenant desires reassuring dependability from the virtualized network fabric – bandwidth guarantee with work conservation, bounded tail latency and resilient reachability. However, the slow convergence of prior works under network dynamics and uncertainties can hardly provide the dependability for tenants. Further, state-of-the-art load balance schemes are guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works. In this paper, we propose vFab, a dependable virtualized fabric framework which can (1) quickly detect network failure in data plane, (2) explicitly select proper paths for all flows, and (3) converge to ideal bandwidth allocation at sub-millisecond. The core idea of vFab is to leverage the programmable data plane to build a fusion of an active edge (e.g., NIC) and an informative core (e.g., switch), where the core sends link status and tenant information to the edge via telemetry to help the latter make a timely and accurate decision on path selection and traffic admission. We fully implement vFab with commodity SmartNICs and programmable switches. Extensive evaluations show that vFab can keep bandwidth guarantee with high bandwidth utilization, low and bounded latency, and resilient reachability under various network scenarios with limited overhead. Application-level experiments show that vFab can improve QPS by$2.4\times $and cut tail latency by$10\times $compared to the alternatives.
Kaihui Gao, Shuai Wang 0028, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Tao Sun 0010
IEEE/ACM Trans. Netw.4
2022 Bandwidth-efficient Microburst Measurement in Large-scale Datacenter Networks
abstract
Microburst measurement is essential for diagnosing and mitigating performance problems in datacenter networks. The key is to efficiently identify the flows that contribute the most to queue buildup. However, because the existing microburst measurement systems capture packet-level information, they incur significant bandwidth overhead. We present BurstScope, a bandwidth-efficient microburst measurement system that can profile the microburst characteristics and the contributing flows. BurstScope detects the microburst-involved packets in egress pipeline, then aggregates the measurement granularity from packet level to flow level by an invertible sketch. Finally, by carefully partitioning the measurement and statistic tasks between the data and control plane, we generate only one telemetry packet for each microburst. We have implemented BurstScope on Barefoot Tofino switches. Testbed-based evaluations show that BurstScope keeps low bandwidth overhead (< 0.02%) and high identification accuracy (> 97%). Compared with the state-of-the-art system, BurstScope can reduce 60 × bandwidth overhead.
Kaihui Gao, Dan Li 0001, Shuai Wang 0028
APNet2
2022 EndGraph: An Efficient Distributed Graph Preprocessing System
abstract
Graph processing mainly includes two stages, namely, preprocessing and algorithm execution. Most previous proposals for performance enhancement of graph processing systems focus on the algorithm execution stage, and simple ignore the preprocessing overhead. However, in this work, we argue that the cost of preprocessing can not be ignored since the preprocessing time is much longer than the algorithm execution time in state-of-the-art systems.We propose EndGraph, a distributed graph preprocessing system, to improve preprocessing performance. Firstly, for graph partitioning, we find existing systems either assign imbalanced preprocessing workloads or spend too much time on graph partitioning. Hence, EndGraph proposes a novel chunk-based partition algorithm to balance preprocessing workloads and achieve theoretical lower bound of time complexity. Secondly, for graph construction (converting data layout from edge array to adjacency list), existing systems use counting sort, which is not efficient for computation and communication. EndGraph employs a novel two-level graph construction method by carefully decoupling the graph construction into intra-machine and inter-machine construction. Our extensive evaluation results show that, compared with five state-of-the-art systems, LFGraph, PowerLyra, PowerGraph, D-Galois, and Gemini, EndGraph can improve the preprocessing performance up to 35.76 ×(from 4.72×). To show the generality of EndGraph, we integrate it with D-Galois and Gemini, and it improves the end-to-end (including preprocessing and algorithm execution) graph processing performance up to 7.44× (from 2.96×).
Tianfeng Liu, Dan Li 0001
ICDCS2
2022 Break the Blackbox! Desensitize Intra-domain Information for Inter-domain Routing
abstract
Along with the ever-increasing amount of data generated from edge networks, cross domain (also known as Autonomous Systems, AS) transmission problem has attracted more and more attention. As mature and widely used inter-domain routing protocols, BGP-based solutions often use the number of domains (i.e. AS hops) of each path to make inter-domain routing decisions, which is simple and effective, but usually can not get the optimal routing results due to the lack of real state/information within ASes. These protocols choose the path with less AS hops as the forwarding path, even if the total latency or cost of the domains on this path is higher. While to solve this problem, directly access to intra-domain information as the assistance to make routing decisions is impractical due to data privacy.In this paper, we propose DIT, which makes near-optimal inter-domain routing decisions with desensitized intra-domain information. To do so, we design a homomorphic encrypted-based private number comparison scheme to export intra-domain information securely and thus assist in routing decisions. We conduct a series of experiments according to five real network topologies with nearly 900 simulated flows, and the results show that DIT reduces the number of forwarding hops by about 45% in average and reduces flow completion time by about 60%.
Peizhuang Cong, Yuchao Zhang 0004, Wendong Wang 0003, Xiangyang Gong, Tong Yang 0003, Dan Li 0001, Ke Xu 0002
IWQoS8
2022 Buffer-based End-to-end Request Event Monitoring in the Cloud
Kaihui Gao, Chen Sun 0005, Shuai Wang 0028, Dan Li 0001, Yu Zhou 0008, Hongqiang Harry Liu, Lingjun Zhu, Ming Zhang 0005
NSDI4
2022 Elixir: A High-performance and Low-cost Approach to Managing Hardware/Software Hybrid Flow Tables Considering Flow Burstiness
Yanshu Wang, Dan Li 0001, Yuanwei Lu
NSDI2
2022 Predictable vFabric on informative data plane
abstract
In multi-tenant data centers, each tenant desires reassuring predictability from the virtual network fabric - bandwidth guarantee, work conservation, and bounded tail latency. Achieving these goals simultaneously relies on rapid and precise traffic admission. However, the slow convergence (tens of milliseconds) of prior works can hardly satisfy the increasingly rigorous performance demand under dynamic traffic patterns. Further, state-of-the-art load balance schemes are all guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works.
Shuai Wang 0028, Kaihui Gao, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Ming Zhang 0005
SIGCOMM4
2022 Themis: Accelerating the Detection of Route Origin Hijacking by Distinguishing Legitimate and Illegitimate MOAS
Lancheng Qin, Dan Li 0001
USENIX Security Symposium2
2022 Impact of Synchronization Topology on DML Performance: Both Logical Topology and Physical Topology
abstract
To tackle the increasingly larger training data and models, researchers and engineers resort to multiple servers in a data center for distributed machine learning (DML). On one hand, DML enables us to leverage the computation power of multiple servers, which can effectively accelerate those computation-intensive tasks. On the other hand, DML also incurs significant communication cost due to parameter synchronization among these servers. In this paper, we want to explore the impact of synchronization topology, including both logical topology and physical topology, on the DML performance. First, we revisit the existing logical topologies, e.g., parameter server and ring allreduce, for parameter synchronization, and we find that theseflatsynchronization topologies is inefficient when running a large-scale DML training. Therefore, we propose a hierarchical parameter synchronization topology, called HiPS, which can achieve efficient parameter synchronization even on a large scale. Then, we compare two representative physical network topologies, namely, Fat-Tree and BCube. Based on our analyses, BCube has many advantages over Fat-Tree, e.g., higher bandwidth, better load balance, and lower hardware cost. The simulation results also show that BCube is more friendly to RDMA. Relying on the advantages of HiPS and BCube, the GST of “HiPS+BCube” is 12% ~ 70% lower than other combinations. Moreover, when the cluster size increases from 16 to 1024, the performance of “HiPS+BCube” only drops by 6.5%, while the performance of “Ring+BCube” drops by 44.6%. Hence, we believe “HiPS+BCube” is the optimal solution to benefit DML in large scale.
Shuai Wang 0028, Jinkun Geng, Dan Li 0001
IEEE/ACM Trans. Netw.3
2021 Mu: An Efficient, Fair and Responsive Serverless Framework for Resource-Constrained Edge Clouds
abstract
Serverless computing platforms simplify development, deployment, and automated management of modular software functions. However, existing serverless platforms typically assume an over-provisioned cloud, making them a poor fit for Edge Computing environments where resources are scarce. In this paper we propose a redesigned serverless platform that comprehensively tackles the key challenges for serverless functions in a resource constrained Edge Cloud.
Viyom Mittal, Shixiong Qi, Ratnadeep Bhattacharya, Xiaosu Lyu, Sameer G. Kulkarni, Dan Li 0001, Jinho Hwang, K. K. Ramakrishnan, Timothy Wood 0001
SoCC7
2021 Fast and Robust Online Traffic Classification Supporting Unseen Applications
abstract
Online traffic classification is a fundamental toolkit in network management, such as QoS and network security. The speed and generalization ability of online classification are two requirements that need to be satisfied simultaneously. However, existing methods may suffer from generalization degradation on the traffic with unseen applications which are constantly emerging in the network, due to the feature distribution drift (FDD) caused by their non-robust feature engineering approaches. Based on Deep Metric Learning which can restrict the distances between samples explicitly and clustering algorithm which can learn multiple clusters within each category, this paper presents Robot, a fast and robust online traffic classification system. At its core, Robot leverages two building blocks to classify high-speed traffic flows: 1) For fast classification, Fast model classifies traffic and detects FDD samples simultaneously based on only one packet. 2) For robust classification, once FDD samples are detected, a flow collector will be triggered to collect flows and then Robust model, a multi-center model, will further identify them based on the hybrid of packet-level and flow-level features. Our comprehensive experiments demonstrate that Robot can achieve comparative classification speed and better generalization ability on the mixed traffic datasets with seen and unseen applications (with the FDD detection accuracy of up to 85.8% and with nearly 10% improvement in classification accuracy), compared with the state-of-the-art methods.
Dan Li 0001, Kaihui Gao
GLOBECOM2
2021 FastKeeper: A Fast Algorithm for Identifying Top-k Real-time Large Flows
abstract
Precise identification of large flows is a critical task in network traffic measurement. Previous works focus on identification of elephant flows (i.e. large flows from the beginning of the measurement). However, we generally observe that the flow rates change periodically and abruptly. In addition, the large flows may become small flows over time. Thus, elephant flows are not equal to the real-time large flows, and previous works cannot be used for the identification of the real-time large flows that is more meaningful for modern network applications. Nevertheless, identification of real-time large flows is challenging in that it requires accurate measurement of real-time flow rates and timely replacement of flows that have become small in the measurement data structure. In this paper, we propose FastKeeper to identify real-time large flows with a primary goal of simultaneously achieving low over-head, high performance and high accuracy. FastKeeper employs a sliding-window-based algorithm for accurate measurement of real-time flow rates and a bitmap-voting algorithm for timely replacement of flows that have become small in the measurement data structure. We evaluate it on DPDK using the traces from an operator network, and the evaluation demonstrates that it achieves high accuracy (98%) and processing throughput (25.33Mpps).
Yanshu Wang, Dan Li 0001
GLOBECOM2
2021 Analyzing Open-Source Serverless Platforms: Characteristics and Performance (S)
abstract
Serverless computing is increasingly popular because of its lower cost and easier deployment.Several cloud service providers (CSPs) offer serverless computing on their public clouds, but it may bring the vendor lock-in risk.To avoid this limitation, many open-source serverless platforms come out to allow developers to freely deploy and manage functions on self-hosted clouds.However, building effective functions requires much expertise and thorough comprehension of platform frameworks and features that affect performance.It is a challenge for a service developer to differentiate and select the appropriate serverless platform for different demands and scenarios.Thus, we elaborate the frameworks and event processing models of four popular open-source serverless platforms and identify their salient idiosyncrasies.We analyze the root causes of performance differences between different service exporting and auto-scaling modes on those platforms.Further, we provide several insights for future work, such as auto-scaling and metric collection.Index Terms-cloud computing,
Sameer G. Kulkarni, K. K. Ramakrishnan, Dan Li 0001
SEKE4
2021 Sphinx: A transport protocol for high-speed and lossy mobile networks
Dan Li 0001, Wenfei Wu, K. K. Ramakrishnan, Jinkun Geng, Fanzhao Wang, Kai Zheng 0003
Comput. Networks2
2021 Accelerating End-to-End Deep Learning Workflow With Codesign of Data Preprocessing and Scheduling
abstract
In this article, we investigate the performance bottleneck of existing deep learning (DL) systems and propose DLBooster to improve the running efficiency of deploying DL applications on GPU clusters. At its core, DLBooster leverages two-level optimizations to boost the end-to-end DL workflow. On the one hand, DLBooster selectively offloads some key decoding workloads to FPGAs to provide high-performance online data preprocessing services to the computing engine. On the other hand, DLBooster reorganizes the computational workloads of training neural networks with the backpropagation algorithm and schedules them according to their dependencies to improve the utilization of GPUs at runtime. Based on our experiments, we demonstrate that compared with baselines, DLBooster can improve the image processing throughput by 1.4× - 2.5× and reduce the processing latency by 1/3 in several real-world DL applications and datasets. Moreover, DLBooster consumes less than 1 CPU core to manage FPGA devices at runtime, which is at least 90 percent less than the baselines in some cases. DLBooster shows its potential to accelerate DL workflows in the cloud.
Dan Li 0001, Binyao Jiang, Jinkun Geng, Wei Bai 0001, Yongqiang Xiong
IEEE Trans. Parallel Distributed Syst.2
2020 CEFS: compute-efficient flow scheduling for iterative synchronous applications
abstract
Iterative Synchronous Applications (ISApps) are popular in today's data centers, represented by distributed deep learning (DL) training. In ISApps, multiple nodes carry out the computing task iteratively, with globally synchronizing the results in each iteration. To increase the scaling efficiency of ISApps, in this paper we propose a new flow scheduling approach, called CEFS. CEFS saves the waiting time of computing nodes from two aspects. For a single node, flows with data which can trigger earlier computation at the node are assigned with higher priority; among nodes, flows towards slower nodes are assigned with higher priority.
Shuai Wang 0028, Dan Li 0001, Jiansong Zhang 0001, Wei Lin 0016
CoNEXT2
2020 Fela: Incorporating Flexible Parallelism and Elastic Tuning to Accelerate Large-Scale DML
abstract
Distributed machine learning (DML) has become the common practice in industry, because of the explosive volume of training data and the growing complexity of training model. Traditional DML follows data parallelism but causes significant communication cost, due to the huge amount of parameter transmission. The recently emerging model-parallel solutions can reduce the communication workload, but leads to load imbalance and serious straggler problems. More importantly, the existing solutions, either data-parallel or model-parallel, ignore the nature of flexible parallelism for most DML tasks, thus failing to fully exploit the GPU computation power. Targeting at these existing drawbacks, we propose Fela, which incorporates both flexible parallelism and elastic tuning mechanism to accelerate DML. In order to fully leverage GPU power and reduce communication cost, Fela adopts hybrid parallelism and uses flexible parallel degrees to train different parts of the model. Meanwhile, Fela designs token-based scheduling policy to elastically tune the workload among different workers, thus mitigating the straggler effect and achieve better load balance. Our comparative experiments show that Fela can significantly improve the training throughput and outperforms the three main baselines (i.e. dataparallel, model-parallel, and hybrid-parallel) by up to 3.23×, 12.22×, and 1.85× respectively.
Jinkun Geng, Dan Li 0001, Shuai Wang 0028
ICDE2
2020 Geryon: Accelerating Distributed CNN Training by Network-Level Flow Scheduling
abstract
Increasingly rich data sets and complicated models make distributed machine learning more and more important. However, the cost of extensive and frequent parameter synchronizations can easily diminish the benefits of distributed training across multiple machines. In this paper, we present Geryon, a network-level flow scheduling scheme to accelerate distributed Convolutional Neural Network (CNN) training. Geryon leverages multiple flows with different priorities to transfer parameters of different urgency levels, which can naturally coordinate multiple parameter servers and prioritize the urgent parameter transfers in the entire network fabric. Geryon requires no modification in CNN models and does not affect the training accuracy. Based on the experimental results of four representative CNN models on a testbed of 8 GPU servers, Geryon achieves up to 95.7% scaling efficiency even with 10GbE bandwidth. In contrast, for most models, the scaling efficiency of vanilla TensorFlow is no more than 37% and that of TensorFlow with parameter partition and slicing is around 80%. In terms of training throughput, Geryon enhanced with parameter partition and slicing achieves up to 4.37x speedup, where the flow scheduling algorithm itself achieves up to 1.2x speedup over parameter partition and slicing.
Shuai Wang 0028, Dan Li 0001, Jinkun Geng
INFOCOM2
2020 Incorporating Intra-flow Dependencies and Inter-flow Correlations for Traffic Matrix Prediction
abstract
Traffic matrix (TM) prediction is essential for effective traffic engineering and network management. Based on our analysis of real traffic traces from Wide Area Network, the traffic flows in TM are both time-varying (i.e. with intra-flow dependencies) and correlated with each other (i.e. with inter-flow correlations). However, most existing works in TM prediction ignore inter-flow correlations. In this paper, we propose a novel Attention-based Convolutional Recurrent Neural Network (ACRNN) model to capture both intra-flow dependencies and inter-flow correlations. ACRNN mainly contains two components: 1) Correlational Modeling employs attention-based convolutional structures to capture the correlation of any two flows in TMs; 2) Temporal Modeling uses attention-based recurrent structures to model the long-term temporal dependencies of each flow, and then predicts TMs according inter-flow correlations and intra-flow dependencies. Experiments on two real-world datasets show that, when predicting the next TM, ACRNN model reduces the Mean Squared Error by up to 44.8% and reduces the Mean Absolute Error by up to 30.6%, compared to state-of-the-art method; and the gap is even larger when predicting the next multiple TMs. Besides, simulation results demonstrate that ACRNN's accurate prediction can help traffic engineering to mitigate traffic congestion.
Kaihui Gao, Dan Li 0001, Li Chen 0008, Jinkun Geng, Fei Gui
IWQoS2
2020 Managing Multicast Membership for Software Defined Data Center Network
abstract
In this paper we design DCMA, a novel multicast membership management scheme for data center networks. Unlike traditional protocols like IGMP/MLD, DCMA leverages the characteristics of data center multicast application and the emerging software defined networking (SDN) technique to manage multicast members in an easier and better way. Multicast application master delivers the membership to the SDN controller, and the SDN controller assigns the group addresses in a coordinated way to minimize the forwarding table size in switches. In particular, by formulating the multicast group address allocation problem and capturing the relationship among forwarding entries in different switches, we design both a batch algorithm and an incremental algorithm to allocate the multicast group addresses. Evaluations based on real-world traces show that, DCMA can save almost half multicast forwarding entries in switches compared with random multicast address allocation.
Dan Li 0001, Jing Zhu 0007, Hongnan Liu, Kai Chen 0005
VTC Fall2
2020 A Scalable, High-Performance, and Fault-Tolerant Network Architecture for Distributed Machine Learning
abstract
In large-scale distributed machine learning (DML), the network performance between machines significantly impacts the speed of iterative training. In this paper we propose BML, a scalable, high-performance and fault-tolerant DML network architecture on top of Ethernet and commodity devices. BML builds on BCube topology, and runs a fully-distributed gradient synchronization algorithm. Compared to a Fat-Tree network with the same size, a BML network is expected to take much less time for gradient synchronization, for both low theoretical synchronization time and its benefit to RDMA transport. With server/link failures, the performance of BML degrades in a graceful way. Experiments of MNIST and VGG-19 benchmarks on a testbed with 9 dual-GPU servers show that, BML reduces the job completion time of DML training by up to 56.4% compared with Fat-Tree running state-of-the-art gradient synchronization algorithm.
Dan Li 0001, Jinkun Geng, Yanshu Wang, Shuai Wang 0028, Shutao Xia
IEEE/ACM Trans. Netw.2
2019 Accelerating Distributed Machine Learning by Smart Parameter Server
abstract
Parameter Server (PS)-based architecture is widely applied in distributed machine learning (DML), but it is still an open issue how to improve the DML performance in this frame-work. Existing works mainly focus on the view of workers. In this paper, we tackle this problem from another perspective, by leveraging the central control on the PS. Specifically, we propose SmartPS, which transforms the passive role of PS in traditional DML and fully exploits the intelligence of PS. Firstly, the PS holds the global view of parameter dependency, facilitating it to update workers' parameters selectively and proactively. Secondly, the PS records the workers' speeds, and prioritizes parameter transmission to narrow the gap between stragglers and fast workers. Thirdly, the PS considers the parameter dependency in consecutive training iterations, and opportunistically blocks unnecessary pushes from workers. We conduct comparative experiments with two typical benchmarks, Matrix Factorization (MF) and PageRank (PR). The experimental results prove that, compared with all the baseline algorithms (i.e. standard BSP, ASP and SSP), SmartPS can reduce the overall training time by 65.7%~84.9%, with the same training accuracy.
Jinkun Geng, Dan Li 0001, Shuai Wang 0028
APNet2
2019 Rima: An RDMA-Accelerated Model-Parallelized Solution to Large-Scale Matrix Factorization
abstract
Matrix factorization (MF) is a fundamental technique in machine learning and data mining, which gains wide application in many fields. When the matrix becomes large, MF cannot be processed on a single machine. Considering this, many distributed SGD algorithms (e.g. DSGD) have been developed to solve large-scale MF on multiple machines in a model-parallel way. Existing distributed algorithms are primarily implemented under Map/Reduce or PS (parameter server)-based architectures, which incur significant communication overheads. Besides, existing solutions cannot well embrace the benefit of RDMA/RoCE transport and suffer from scalability problems. Targeting at these drawbacks, we propose Rima, which uses ring-based model parallelism to solve large-scale MF with higher communication efficiency. Compared with PS-based SGD algorithms, Rima also consumes less queue pairs (QPs) and can thus better leverage the power of RDMA/RoCE to accelerate the training speed. Our experiment shows that, compared with PS-based DSGD when solving 1M × 1M MF, Rima achieves comparable convergence performance after equal number of iterations, but reduces the training time by 68.7% and 85.4% via TCP and RDMA respectively.
Jinkun Geng, Dan Li 0001, Shuai Wang 0028
ICDE2
2019 DLBooster: Boosting End-to-End Deep Learning Workflows with Offloading Data Preprocessing Pipelines
abstract
In recent years, deep learning (DL) has prospered again due to improvements in both computing and learning theory. Emerging studies mostly focus on the acceleration of refining DL models but ignore data preprocessing issues. However, data preprocessing can significantly affect the overall performance of end-to-end DL workflows. Our studies on several image DL workloads show that existing preprocessing backends are quite inefficient: they either perform poorly in throughput (30% degradation) or burn too many (>10) CPU cores. Based on these observations, we propose DLBooster, a high-performance data preprocessing pipeline that selectively offloads key workloads to FPGAs, to fit the stringent demands on data preprocessing for cutting-edge DL applications. Our testbed experiments show that, compared with the existing baselines, DLBooster can achieve 1.35×~2.4× image processing throughput in several DL workloads, but consumes only 1/10 CPU cores. Besides, it also reduces the latency by 1/3 in online image inference.
Dan Li 0001, Binyao Jiang, Xi Fan, Jinkun Geng, Wei Bai 0001, Ran Shu 0001, Peng Cheng 0005, Yongqiang Xiong
ICPP2
2019 Impact of Network Topology on the Performance of DML: Theoretical Analysis and Practical Factors
abstract
To deal with the increasingly larger input data and model sizes, it has become necessary to scale the training of machine learning models to multiple nodes, even a server cluster, which we call distributed machine learning, or DML. However, DML utilizes more computation power at the cost of high communication overhead, which may limit the overall performance in turn. In this paper, we study the impact of network topology on the DML performance both in theory and in practice. We compare two representative network topologies, namely, Fat-Tree which is widely-used in modern data centers, and BCube, which is a low-cost and server-centric network topology, both running on top of RDMA. The results show that Fat-Tree not only has theoretically higher global synchronization time (GST) than BCube, but its practical GST (by NS-3 based simulation) is also considerably larger than the theoretical one. By analyzing the large-scale simulation traces, we find that the root cause for the gap in Fat-Tree comes from the load imbalance among the multiple parallel paths as well as the inevitable PFC frames, both of which do not appear in BCube. For a cluster of around 250 servers, BCube achieves 53%\sim 70% lower GST than Fat-Tree from the simulation. As a result, we suggest using server-centric network topology such as BCube, instead of the common Fat-Tree network, to build a special-purpose DML cluster, due to its parallel synchronization, RDMA friendliness, natural load balance, as well as low economical cost.
Shuai Wang 0028, Dan Li 0001, Jinkun Geng
INFOCOM2
2019 Sphinx: A Transport Protocol for High-Speed and Lossy Mobile Networks
abstract
Modern mobile wireless networks have been demonstrated to be high-speed but lossy, while mobile applications have more strict requirements including reliability, goodput guarantee, bandwidth efficiency, and computation efficiency. Such a complicated combination of requirements and conditions in networks pushes the pressure to transport layer protocol design. We analyze and argue that few of existing network transport layer solutions are able to handle all these requirements. We design and implement Sphinx to satisfy the four requirements in high-speed and lossy networks. Sphinx has (1) a proactive coding-based method named semi-random LT codes for loss recovery, which estimates packet loss rate and adjusts the redundancy level accordingly, (2) a reactive retransmission method named Instantaneous Compensation Mechanism (ICM) for loss retransmission, which compensates the lost packets once actual loss exceeds the estimation, and (3) a parallel coding architecture, which leverages multi-core, shared memory and kernel-bypass DPDK. Prototype and evaluation show that Sphinx outperforms TCP and other coding solutions significantly in microbenchmarks across all four requirements, and improves the performance of applications such as video streaming and block data transfer.
Dan Li 0001, Wenfei Wu, K. K. Ramakrishnan, Jinkun Geng, Fei Gui, Fanzhao Wang, Kai Zheng 0003
IPCCC2
2019 Metro: An Efficient Traffic Fast Rerouting Scheme With Low Overhead
abstract
Failure is common instead of exception in large-scale networks. To provide high service quality to upper-layer applications, it is desired that a converged backup path can be rapidly launched when failure occurs. In this paper, we design an IP based Fast ReRouting (FRR) scheme called Metro, which can solve the traffic rerouting convergence problem after arbitrary single link/node failure with low stretch for the backup path. When failure occurs in the network, Metro first indicates all the network areas that would be affected by the failure, and then finds out a few bridge links to drain the traffic in the affected network area to the network area that is not affected by the failure. In this way, Metro does not configure tunnels, encapsulate or modify data packets, and hence it is easy to be deployed in current networks. Extensive simulations show that Metro can solve arbitrary single link/node failure with backup paths shorter than the state-of-the-art solutions, and about 98% of the backup path stretch in Metro are the same as the optimal tunnel scheme.
Xuya Jia, Dan Li 0001, Jing Zhu 0007, Yong Jiang 0001
IEEE/ACM Trans. Netw.2
2018 Dante: Enabling FOV-Aware Adaptive FEC Coding for 360-Degree Video Streaming
abstract
As 360-degree videos grow dramatically in popularity, more applications demand the ability to stream 360-degree videos to wirelessly connected devices, such as smartphone headsets. However, the limited capacity and the unstable network conditions make wireless networks ill-suited to the requirements of 360-degree videos--high resolution and low delay. One common approach is to take advantage of the fact that the viewer only watches a small portion of the video around the field of view (FOV). This allows for better allocation of network bandwidth by prioritizing content the viewer actually watches. Previous efforts on 360-degree videos have largely focused on adapting the encoded bitrate to optimize video quality in the time-varying FOV. This paper follows the general FOV-aware approach but uses a different technique. Rather than adapting bitrate, we explore the opportunities of a custom underlying transport protocol for 360-degree videos. In particular, we make a case for using Forward Error Correction (FEC) coding over UDP to reduce video streaming delay (a key limitation of all TCP-based approaches). We present Dante, an FOV-aware UDP-based video streaming protocol that adapts to changing network conditions by dynamically choosing FEC redundancy levels based on how close the video content is to the FOV region. Experimental results show that Dante improves video quality (PSNR) by 20% to 30% over traditional UDP-based video streaming protocols and 40% over FOV-aware DASH.
Zhetao Li, Fei Gui, Jinkun Geng, Dan Li 0001, Zhibo Wang 0001, Usama Zafar
APNet4
2018 BML: A High-performance, Low-cost Gradient Synchronization Algorithm for DML Training
abstract
In distributed machine learning (DML), the network performance between machines significantly impacts the speed of iterative training. In this paper we propose BML, a new gradient synchronization algorithm with higher network performance and lower network cost than the current practice. BML runs on BCube network, instead of using the traditional Fat-Tree topology. BML algorithm is designed in such a way that, compared to the parameter server (PS) algorithm on a Fat-Tree network connecting the same number of server machines, BML achieves theoretically 1/k of the gradient synchronization time, with k/5 of switches (the typical number of k is 2∼4). Experiments of LeNet-5 and VGG-19 benchmarks on a testbed with 9 dual-GPU servers show that, BML reduces the job completion time of DML training by up to 56.4%.
Dan Li 0001, Jinkun Geng, Yanshu Wang, Shuai Wang 0028, Shutao Xia
NeurIPS2
2018 Towards full virtualization of SDN infrastructure
Dan Li 0001, Yirong Yu, Jing Zhu 0007, Jinkun Geng
Comput. Networks2
2018 SVDC: A Highly Scalable Isolation Architecture for Virtualized Layer-2 Data Center Networks
abstract
While large layer-2 networks are widely accepted as the network fabric for modern data centers and network virtualization is required to support multi-tenant cloud computing, existing network virtualization solutions are not specifically designed for layer-2 networks. In this paper, we designSVDC, a highly-scalable and low-overhead virtualization architecture for large layer-2 data center networks. By leveraging the emerging software defined networking (SDN) framework, SVDC decouples the global identifier of a virtual network from the identifier carried in the packet header. Hence, SVDC can scale to a great number of virtual networks with a very short tag in the packet header, which is never achieved by previous network virtualization solutions. SVDC enhances MAC-in-MAC encapsulation in a way that packets with overlapped MAC addresses are correctly forwarded even without in-packet global identifiers to differentiate the virtual networks they belong to. Besides, scalable and efficient layer-2 multicast and broadcast within virtual networks are also supported in SVDC. With extensive simulations and experiments, we show that SVDC is better than existing solutions in many aspects, particularly isolating virtual networks with high scalability and higher network goodput due to minimal packet header overhead.
Congjie Chen, Dan Li 0001, Jun Li 0001, Konglin Zhu
IEEE Trans. Cloud Comput.2
2018 Dependency-Aware Data Locality for MapReduce
abstract
MapReduce effectively partitions and distributes computation workloads to a cluster of servers, facilitating today's big data processing. Given the massive data to be dispatched, and the intermediate results to be collected and aggregated, there have been a significant studies on data locality that seeks to co-locate computation with data, so as to reduce cross-server traffic in MapReduce. They generally assume that the input data have little dependency with each other, which however is not necessarily true for that of many real-world applications, and we show strong evidence that the finishing time of MapReduce tasks can be greatly prolonged with such data dependency. In this paper, we present Dependency-Aware Locality for MapReduce (DALM) for processing the real-world input data that can be highly skewed and dependent. DALM accommodates data-dependency in a data-locality framework, organically synthesizing the key components from data reorganization, replication, placement. Beside algorithmic design within the framework, we have also closely examined the deployment challenges, particularly in public virtualized cloud environments, and have implemented DALM on Hadoop 1.2.1 with Giraph 1.0.0. Its performance has been evaluated through both simulations and real-world experiments, and compared with that of state-of-the-art solutions.
Xiaoqiang Ma, Xiaoyi Fan 0001, Jiangchuan Liu, Dan Li 0001
IEEE Trans. Cloud Comput.4
2017 LOS: A High Performance and Compatible User-level Network Operating System
abstract
With the ever growing speed of Ethernet NIC and more and more CPU cores on commodity X86 servers, the processing capability of the network stack in Linux kernel has become the bottleneck. Recently there is a trend on moving the network stack up to user level and bypassing the kernel. However, most of these stacks require changing the APIs or modifying the source code of applications, and hence are difficult to support legacy applications. In this work, we design and develop LOS, a user-level network operating system that not only gains high throughput and low latency by kernel-bypass technologies but also achieves compatibility with legacy applications. We successfully run Nginx and NetPIPE on top of LOS without touching the source code, and the experimental results show that LOS achieves significant throughput and latency gains compared with Linux kernel.
Jinkun Geng, Du Lin, Ruilin Ling, Dan Li 0001
APNet7
2017 Survivable and bandwidth-guaranteed embedding of virtual clusters in cloud data centers
abstract
Cloud computing has emerged as a powerful and elastic platform for internet service hosting, yet it also draws concerns of the unpredictable performance of cloud-based services due to network congestion. To offer predictable performance, the virtual cluster abstraction of cloud services has been proposed, which enables allocation and performance isolation regarding both computing resources and network bandwidth in a simplified virtual network model. One issue arisen in virtual cluster allocation is the survivability of tenant services against physical failures. Existing works have studied virtual cluster backup provisioning with fixed primary embeddings, but have not considered the impact of primary embeddings on backup resource consumption. To address this issue, in this paper we study how to embed virtual clusters survivably in the cloud data center, by jointly optimizing primary and backup embeddings of the virtual clusters. We formally define the survivable virtual cluster embedding problem. We then propose a novel algorithm, which computes the most resource-efficient embedding given a tenant request. Since the optimal algorithm has high time complexity, we further propose a faster heuristic algorithm, which is several orders faster than the optimal solution, yet able to achieve similar performance. Besides theoretical analysis, we evaluate our algorithms via extensive simulations.
Ruozhou Yu, Guoliang Xue, Xiang Zhang 0005, Dan Li 0001
INFOCOM4
2017 Quick NAT: High performance NAT system on commodity platforms
abstract
NAT gateway is an important network system in today's IPv4 network when translating a private IPv4 address to a public address. However, traditional NAT system based on Linux Netfilter cannot achieve high network throughput to meet modern requirements such as data centers. To address this challenge, we improve the network performance of NAT system by three ways. First, we leverage DPDK to enable polling and zero-copy delivery, so as to reduce the cost of interrupt and packet copies. Second, we enable multiple CPU cores to process in parallel and use lock-free hash table to minimize the contention between CPU cores. Third, we use hash search instead of sequential search when looking up the NAT rule table. Evaluation shows that our Quick NAT system significantly improves the performance of NAT on commodity platforms.
Dan Li 0001, Ruilin Ling
LANMAN2
2017 A survey of network update in SDN
Dan Li 0001, Konglin Zhu, Shutao Xia
Frontiers Comput. Sci.1
2017 Network Performance Aware Optimizations on IaaS Clouds
abstract
Network performance aware optimizations have long been a hot research topic to optimize distributed applications on traditional network environments. However, those optimization techniques rely on a few measurements on pair-wise network performance, and such direct use of network measurements is no longer valid on Infrastructure-as-a-service (IaaS) clouds. First, the direct calibration is ineffective. Network performance measurements may not represent the long-term performance (informally the stable component inside network performance) because of virtualization and network performance interference in the cloud. Second, the direct calibration is inefficient because the measurement overhead of all pair-wise link performance in a cluster becomes prohibitively high as the number of instances increases. To effectively and efficiently utilize existing network performance aware optimizations on IaaS clouds, we propose to reduce the measurement overhead and decouple the constant component from the dynamic network performance while minimizing the difference between the network performance and the constant component. For effectiveness, we use the constant component to guide the network performance aware optimizations. For efficiency, we exploit a non-negative matrix factorization (NMF) method to reduce the calibration overhead. Furthermore, we observe a tradeoff between effectiveness and efficiency, and develop an adaptive approach to capture this tradeoff. We demonstrate effectiveness and efficiency of our approach by adopting network performance aware optimizations on two kinds of basic applications, collective communications of MPI and generic topology mapping, and two real-world applications, namely N-body and conjugate gradient (CG). Our experiments on Amazon EC2 and simulations demonstrate significant calibration overhead reduction and performance improvement on guiding network performance aware optimizations, when comparing our approach to other state-of-the-art approaches.
Yifan Gong 0003, Bingsheng He, Dan Li 0001
IEEE Trans. Computers3
2017 𝔽2 Tree: Rapid Failure Recovery for Routing in Production Data Center Networks
abstract
Failures are not uncommon in production data center networks (DCNs) nowadays. It takes long time for the DCN routing to recover from a failure and find new forwarding paths, significantly impacting realtime and interactive applications at the upper layer. In this paper, we present a fault-tolerant DCN solution, called F2Tree, which is readily deployed in existing DNCs. F2Tree can significantly improve the failure recovery time only through a small amount of link rewiring and switch configuration changes. Through testbed and emulation experiments, we show that F2Tree can greatly reduce the routing recovery time after failure (by 78%) and improve the performance of upper layer applications when routing failure happens (96% less deadline-missing requests).
Guo Chen 0001, Youjian Zhao, Hailiang Xu, Dan Pei, Dan Li 0001
IEEE/ACM Trans. Netw.5
2016 DVMP: Incremental traffic-aware VM placement on heterogeneous servers in data centers
abstract
As the tremendous momentum cloud computing has grown, the modern data center networks are facing challenge to handle the increasing traffic demand among virtual machines (VMs). Simply adding more switches and links may increase network capacity but at the same time increase the complexity and infrastructure cost. Thus, intelligent VM placement has been proposed to reduce the intra-DC traffic. Prior solutions model the traffic-aware VM placement problem as a Balanced Minimum K-cut Problem (BMKP). However, the assumptions of “once-for-all” VM placement on physical servers with equal VM slots are often not realistic in practical data centers, and thus the naive BMKP model may lead to suboptimal placement solutions. In this work, we revisit the problem by considering the server heterogeneity and propose an incremental traffic-aware VM placement algorithm. Given that the BMKP model cannot be directly applied, we make a number of transformations to re-establish the model. First, by introducing pseudo VM slots on physical servers with less VM slots, we allow the number of available VM slots of each server to be different. Second, pseudo edges with infinite costs are added between existing VMs, and thus previously deployed VMs on the same physical server will still be packed together. Third, a change on the number of pseudo VM slots is applied, so that existing VMs placed on different physical servers will still be separated. In this way, we reduce the problem to a new BMKP problem, which results in a much better solution. The evaluation results show that DVMP can reduce up to 28%, 39% and 55% traffic compared with naive BMKP model, greedy VM placement and random VM placement, respectively.
Dan Li 0001, Syed Shah-e-Mardan Ali Rizvi, Fangxin Wang 0001, Wu He
IWQoS1
2016 PALS: Saving Network Power With Low Overhead to ISPs and Applications
abstract
Power saving in the network infrastructure has received great attention in recent years. Power-aware traffic management is proposed in many works, in which a subset of routers/links are preferentially used to carry traffic while other links are activated only when traffic load is high. However, it remains challenging how to minimize the overhead to both ISPs and applications, which is important to the successful deployment of power-aware traffic management in a real network. This paper presents PALS, a new Power-Aware Link State routing based traffic management protocol. Compared with previous solutions, PALS remarkably reduces the overheads to ISPs and applications by the following innovations. First, PALS minimizes the forwarding table expansion due to dynamic power-aware routing, by using destination based routing instead of pairwise routing (e.g., MPLS). Second, PALS limits packet reordering for applications, by never splitting traffic between an IE (ingress-egress) router pair to multiple paths. Third, PALS significantly reduces the computation complexity of the power-aware routing algorithm, by running a simple path selection algorithm at each ingress router with the knowledge of local traffic information as well as global link utilization, which are much easier to obtain than global traffic matrix required by the state-of-the-art solutions (e.g., Zhang , IEEE ICNP 2010). Extensive simulations and testbed experiments show that, although bearing the simplicities to minimize the overhead, PALS saves satisfactory network power, with quick response to traffic variance and negligible impact on the packet delivery performance for applications.
Dan Li 0001, Yirong Yu, Junxiao Shi, Beichuan Zhang 0001
IEEE/ACM Trans. Netw.1
2016 CCDN: Content-Centric Data Center Networks
abstract
Data center networks continually seek higher network performance to meet the ever increasing application demand. Recently, researchers are exploring the method to enhance the data center network performance by intelligent caching and increasing the access points for hot data chunks. Motivated by this, we come up with a simple yet useful caching mechanism for generic data centers, i.e., a server caches a data chunk after an application on it reads the chunk from the file system, and then uses the cached chunk to serve subsequent chunk requests from nearby servers. To turn the basic idea above into a practical system and address the challenges behind it, we design content-centric data center networks (CCDNs), which exploits an innovative combination of content-based forwarding and location [Internet Protocol (IP)]-based forwarding in switches, to correctly locate the target server for a data chunk on a fully distributed basis. Furthermore, CCDN enhances traditional content-based forwarding to determine the nearest target server, and enhances traditional location (IP)-based forwarding to make high utilization of the precious memory space in switches. Extensive simulations based on real-world workloads and experiments on a test bed built with NetFPGA prototypes show that, even with a small portion of the server's storage as cache (e.g., 3%) and with a modest content forwarding information base size (e.g., 1000 entries) in switches, CCDN can improve the average throughput to get data chunks by 43% compared with a pure Hadoop File System (HDFS) system in a real data center.
Dan Li 0001, Fangxin Wang 0001, Anke Li, K. K. Ramakrishnan, Ying Liu 0024, Xue (Steve) Liu
IEEE/ACM Trans. Netw.2
2016 DCloud: Deadline-Aware Resource Allocation for Cloud Computing Jobs
abstract
With the tremendous growth of cloud computing, it is increasingly critical to provide quantifiable performance to tenants and to improve resource utilization for the cloud provider. Though many recent proposals focus on guaranteeing job performance (with a particular note on network bandwidth) in the cloud, they usually lack efficient utilization of cloud resource, or vice versa. In this paper we present DCloud, which leverages the (soft) deadlines of cloud computing jobs to enable flexible and efficient resource utilization in data centers. With the deadline requirement of a job guaranteed, DCloud employs both time sliding (postponing the launching time of a job) and bandwidth scaling (adjusting the bandwidth associated with VMs) in resource allocation, so as to better match the resource allocated to the job with the cloud's residual resource. Extensive simulations and testbed experiments show that DCloud can accept much more jobs than existing solutions, and significantly increase the cloud provider's revenue with less cost for individual tenants.
Dan Li 0001, Congjie Chen, Junjie Guan, Jing Zhu 0007, Ruozhou Yu
IEEE Trans. Parallel Distributed Syst.1
2015 Rewiring 2 Links Is Enough: Accelerating Failure Recovery in Production Data Center Networks
abstract
Failures are not uncommon in production data center networks (DCNs) nowadays, and it takes long time for the network to recover from a failure and find new forwarding paths, significantly impacting real time and interactive applications at the upper layer. The slow failure recovery is due to two primary reasons. First, there lacks immediate backup paths for downward links in DCN with multi-rooted tree topology. Second, distributed routing protocols in DCN take time to converge after failures. In this paper, we present a fault-tolerant DCN solution, called F2Tree, that can significantly improve the failure recovery time in current DCNs, only through a small amount of link rewiring and switch configuration changes. Because F2Tree does not change any existing software or hardware, it is readily deployed in production DCNs, where other existing proposals fail to achieve. Through testbed and emulation experiments, we show that F2Tree can greatly reduce the time of failure recovery by 78%. Our experimental results also show that, for partition-aggregate applications (popular in DCN) under various failure conditions, F2Tree reduces the ratio of deadline-missing requests by more than 96% compared to current DCNs.
Guo Chen 0001, Youjian Zhao, Dan Pei, Dan Li 0001
ICDCS4
2015 SVirt: A Substrate-agnostic SDN Virtualization Architecture for Multi-tenant Cloud
abstract
Data center operators are accepting software defined networking (SDN) to manage their networks, but it remains challenging how to provide desirable virtual SDN services to tenants in a public cloud. We design SVirt, which enables highly flexible virtual SDN in a multi-tenant cloud by a substrate-agnostic SDN virtualization architecture. By redesigning the physical switch's processing pipeline with a "late-binding key extractor", SVirt supports virtual SDN switches with different processing pipelines simultaneously on a physical switch. In the control plane, SVirt enables "many-to-one" and "one-to-many" mapping when allocating the physical resource for a virtual network, which embraces arbitrary topology and TCAM resource demanded by a virtual network. In the data plane, SVirt explicitly carries the forwarding context information in the packets, overcoming the "context-loss problem" in a virtual SDN network. We develop a NetFPGA prototype of SVirt switch. Evaluations based on event-driven simulations and prototype-based experiments demonstrate that, compared with traditional approaches, SVirt significantly enhances the cloud's capability to accept various virtual SDN requests and improves the network's throughput.
Yirong Yu, Dan Li 0001
ICNP2
2015 TAPS: Software Defined Task-Level Deadline-Aware Preemptive Flow Scheduling in Data Centers
abstract
Many data center applications have deadline requirements, which pose a requirement of deadline-awareness in network transport. Completing within deadlines is a necessary requirement for flows to be completed. Transport protocols in current data centers try to share the network resources fairly and are deadline-agnostic. Recently several works try to address the problem by making as many flows meet deadlines as possible. However, for many data center applications, a task cannot be completed until the last flow finishes, which indicates the bandwidths consumed by completed flows are wasted if some flows in the task cannot meet deadlines. In this paper we design a task-level deadline-aware preemptive flow scheduling(TAPS), which aims to make more tasks meet deadlines. We leverage software defined networking (SDN) technology and generalize SDN from flow-level awareness to task-level awareness. The scheduling algorithm runs on the SDN controller, which decides whether a flow should be accepted or discarded, pre-allocates the transmission time slices and computes the routing paths for accepted flows. Extensive flow-level simulations demonstrate TAPS outperforms Varys, Bara at, PDQ (Preemptive Distributed Quick flow scheduling), D3 (Deadline-Driven Delivery control protocol) and Fair Sharing transport protocols in deadline sensitive data center environment. A simple implementation on real systems also proves that TAPS makes high effective utilization of the network bandwidth in data centers.
Dan Li 0001
ICPP2
2015 MIFO: Multi-path Interdomain Forwarding
abstract
Today's interdomain routing is traffic agnostic when determining the single, best forwarding path. Naturally, as it does not adapt to congestion, the path chosen is not always optimal. In this paper, we focus on designing a multi-path interdomain forwarding (MIFO) mechanism, where AS border routers adaptively forward outbound traffic from a congested default path to an alternative path, without touching the interdomain routing protocols. Different from previous efforts which enable multi-path on control plane, MIFO achieves multi-path on data plane. The multiple alternative forwarding paths are obtained by exploring local BGP RIB. Multi-path forwarding on data plane can create a loop even within a stable network. MIFO solves this problem with a simple and practical approach. Several other challenges are also addressed including preventing cycling packet between iBGP peers and choosing the best alternative path from among multiple candidates. Our evaluations show that MIFO significantly improves the end-to-end throughput at the AS level, compared to traditional BGP and MIRO. For example, with only 50% of the ASes being MIFO capable, a significant percentage of the flows (about 40%) can use at least 50% of the inter-AS link capacity. In contrast, BGP and MIRO routing make less effective use of the inter-AS links, with only 7% and 17% of the flows can be so. Finally, we have developed a prototype implementation of MIFO on Linux with the forwarding engine in the kernel, with the routing daemon developed on XORP platform. The experiments on a test bed built with prototypes show that MIFO can improves the aggregate throughput by 81% compared with BGP routing.
Dan Li 0001, Ying Liu 0024, Dan Pei, K. K. Ramakrishnan
ICPP2
2015 Rapier: Integrating routing and scheduling for coflow-aware data center networks
abstract
In the data flow models of today's data center applications such as MapReduce, Spark and Dryad, multiple flows can comprise a coflow group semantically. Only completing all flows in a coflow is meaningful to an application. To optimize application performance, routing and scheduling must be jointly considered at the level of a coflow rather than individual flows. However, prior solutions have significant limitation: they only consider scheduling, which is insufficient. To this end, we present Rapier, a coflow-aware network optimization framework that seamlessly integrates routing and scheduling for better application performance. Using a small-scale testbed implementation and large-scale simulations, we demonstrate that Rapier significantly reduces the average coflow completion time (CCT) by up to 79.30% compared to the state-of-the-art scheduling-only solution, and it is readily implementable with existing commodity switches.
Yangming Zhao, Kai Chen 0005, Wei Bai 0001, Minlan Yu, Chen Tian 0001, Yanhui Geng, Yiming Zhang 0003, Dan Li 0001, Sheng Wang 0006
INFOCOM8
2015 Bandwidth guaranteed virtual network function placement and scaling in datacenter networks
abstract
Enterprises deploy their middlebox services in cloud seeking for easy management, flexible scalability and economic savings. However, existing elastic virtual network function(VNF) placement strategy often leads to an unpredictable placing location due to the ever-changing workload, which may waste much precious bandwidth resource and bring a lot of VM operation overhead(e.g. VM launch, termination and migration). A key problem for cloud providers is how to conduct an effective service placement and provide resource provision according to various workload, satisfying the bandwidth requirement of each service while saving as much cloud resource as possible. In this paper we solve both the virtual network function(VNF) placement and scaling problem based on preplanned allocation with bandwidth guarantee. We first propose a concept of VNF instance communication graph to describe the bandwidth demand of each VNF instance and explore the placement requirement for bandwidth savings. Then we design an on-line heuristic algorithm to achieve approximate optimal allocation. At last, we also provide an off-line optimal solution for comparison. Our simulation shows that our heuristic solution saves 20% more bandwidth resource and reduce more VM migration overhead than existing elastic placement solution. Its performance is also very close to the optimal solution.
Fangxin Wang 0001, Ruilin Ling, Jing Zhu 0007, Dan Li 0001
IPCCC4
2015 SIONA: A Service and Information Oriented Network Architecture
Mingwei Xu 0001, Zhongxing Ming, Chunmei Xia, Jia Ji, Dan Li 0001, Dan Wang 0002
J. Netw. Comput. Appl.5
2015 On the Network Power Effectiveness of Data Center Architectures
abstract
Cloud computing not only requires high-capacity data center networks to accelerate bandwidth-hungry computations, but also causes considerable power expenses to cloud providers. In recent years many advanced data center network architectures have been proposed to increase the network throughput, such as Fat-Tree [1] and BCube [2], but little attention has been paid to the power efficiency of these network architectures. This paper makes the first comprehensive comparison study for typical data center networks with regard to their Network Power Effectiveness(NPE), which indicates the end-to-end bps per watt in data transmission and reflects the tradeoff between power consumption and network throughput. We take switches, server NICs and server CPU cores into account when evaluating the network power consumption. We measure NPE under both regular routing and power-aware routing, and investigate the impacts of topology size, traffic load, throughput threshold in power-aware routing, network power parameter as well as traffic pattern. The results show that in most cases Flattened Butterfly possesses the highest NPE among the architectures under study, and server-centric architectures usually have higher NPEs than Fat-Tree and VL2 architectures. In addition, the sleep-on-idle technique and power-aware routing can significantly improve the NPEs for all the data center architectures, especially when the traffic load is low. We believe that the results are useful for cloud providers, when they design/upgrade data center networks or employ network power management.
Yunfei Shang, Dan Li 0001, Jing Zhu 0007, Mingwei Xu 0001
IEEE Trans. Computers2
2015 Guaranteeing Heterogeneous Bandwidth Demand in Multitenant Data Center Networks
abstract
The ability to provide guaranteed network bandwidth for tenants is essential to the prosperity of cloud computing platforms, as it is a critical step for offering predictable performance to applications. Despite its importance, it is still an open problem for efficient network bandwidth sharing in a multitenant environment, especially when applications have diverse bandwidth requirements. More precisely, it is not only that different tenants have distinct demands, but also that one tenant may want to assign bandwidth differently across her virtual machines (VMs), i.e., the heterogeneous bandwidth requirements. In this paper, we tackle the problem of VM allocation with bandwidth guarantee in multitenant data center networks. We first propose an online VM allocation algorithm that improves on the accuracy of the existing work. Next, we develop a VM allocation algorithm under heterogeneous bandwidth demands. We conduct extensive simulations to demonstrate the efficiency of our method.
Dan Li 0001, Jing Zhu 0007, Junjie Guan
IEEE/ACM Trans. Netw.1
2015 Willow: Saving Data Center Network Energy for Network-Limited Flows
abstract
Today's giant data centers are power hungry. Data center energy saving not only helps control the operational cost, but also benefits the sustainable growth of cloud services. Due to the adoption of much more switches in modern data centers as well as the mature server-side power management techniques, energy saving for the data center network is becoming increasingly important. Most previous works on saving data center network energy focus on aggregating flows to as few switches as possible. However, in this paper we argue that this method may not work for network-limited flows, the throughputs of which are elastic based on the competing flows. To save the network energy consumed by this kind of elastic flows, we propose a flow scheduling approach called Willow, which takes both the number of switches involved and their active working durations into consideration. We formulate this problem by programming and design a greedy approximate algorithm to schedule flows in an online manner. Simulations based on MapReduce traces show that Willow can save up to 60 percent network energy compared with ECMP scheduling in typical settings, and outperforms other classical heuristic algorithms such as simulated annealing and particle swarm optimization. Testbed Experiments demonstrate that this kind of dynamic energy-efficient flow scheduling causes negligible impact on upper-layer applications.
Dan Li 0001, Yirong Yu, Wu He, Kai Zheng 0003, Bingsheng He
IEEE Trans. Parallel Distributed Syst.1
2014 Dependency-Aware Data Locality for MapReduce
abstract
Recent years have witnessed the prevalence of MapReduce-based systems, e.g., the Apache Hadoop, in large-scale distributed data processing. Fetching data from remote servers across multiple network switches is known to be costly. Hence, it is highly desirable to co-locate computation with data. State-of-the-art popularity-based replication achieves data locality through replicating popular files and spreading the replicas over multiple servers. While working well for independent files, they can store highly dependent files in different servers, resulting in excessive remote data accesses exchanges and consequently prolonging the job completion time. In this paper, we develop DALM (Dependency-Aware Locality for MapReduce), a novel replication strategy for general real-world input data that can be highly skewed and dependent. DALM accommodates data-dependency in a data-locality framework that comprehensively weights such key factors as popularity and storage budget. We extensively evaluate DALM through both simulations and real-world implementations, and have compared with state-of-the-art solutions, including the Hadoop system and the popularity-based Scarlett. The results show that DALM can significantly improve data locality for different inputs. For a popular iterative graph processing application on Hadoop, our prototype implementation of DALM reduces the remote data access and job completion time by 34.3% and 9.4%, respectively.
Xiaoyi Fan 0001, Xiaoqiang Ma, Jiangchuan Liu, Dan Li 0001
IEEE CLOUD4
2014 CDRDN: Content Driven Routing in Datacenter Network
abstract
A major challenge in data center networks is to provide enough network capacity to meet the ever increasing demand of large-scale distributed computing. While existing proposals focus on adding more switches and links, which cost extra hardware and energy, we explore another dimension in which spare disk space at servers is used for caching data, so as to increase network throughput without additional hardware. Leveraging the Named Data Networking (NDN) architecture and unique characteristics of data centers, we design a novel Content Driven Routing in Datacenter Network (CDRDN) that enables universal caching in data center networks in an efficient and scalable way. First, rather than using switches for “on-path” caching like in native NDN, CDRDN uses the large storage space at servers to do “off-path” caching. Second, by taking advantage of data center's regular and hierarchical network topology, CDRDN switches are able to direct requests to nearby server caches even under high dynamics of caches. Third, given the vast amount of data in data centers, the full name-based routing table would be difficult to fit in switch's limited fast memory. CDRDN adopts a compound content and location routing to ensure packet delivery while benefiting from name-based routing as much as the switches can afford. CDRDN extends NDN's adaptive forwarding mechanism to deal with cache misses, link failures, and congestion without running routing protocols or cache exchange protocols. Our packet-level simulations show that CDRDN can almost double the network throughput compared with shortest-path routing under the same setting, and CDRDN can effectively deal with link failures using adaptive forwarding.
Dan Li 0001, Ying Liu 0024
ICCCN2
2014 Freeway: Adaptively Isolating the Elephant and Mice Flows on Different Transmission Paths
abstract
The network resource competition of today' data enters is extremely intense between long-lived elephant flows and latency-sensitive mice flows. Achieving both goals of high throughput and low latency respectively for the two types of flows requires compromise, which recent research has not successfully solved mainly due to the transfer of elephant and mice flows on shared links without any differentiation. However, current data enters usually adopt clos-based topology, e.g. Fat-tree/VL2, so there exist multiple shortest paths between any pair of source and destination. In this paper, we leverage on this observation to propose a flow scheduling scheme, Freeway, to adaptively partition the transmission paths into low latency paths and high throughput paths respectively for the two types of flows. An algorithm is proposed to dynamically adjust the number of the two types of paths according to the real-time traffic. And based on these separated transmission paths, we propose different flow type-specific scheduling and forwarding methods to make full utilization of the bandwidth. Our simulation results show that Freeway significantly reduces the delay of mice flow by 85.8% and achieves 9.2% higher throughput compared with Hedera.
Wei Wang 0157, Yi Sun 0004, Kai Zheng 0003, Mohamed Ali Kâafar, Dan Li 0001, Zhongcheng Li
ICNP5
2014 TED: Inter-domain traffic engineering via deflection
abstract
As inter-domain routing on today's Internet does not and basically cannot consider traffic load when determining best traffic forwarding paths, it is not always optimal for a router to forward packets along its default path, especially when the router's default output port incurs a long queuing delay. In this paper, we design a new approach called TED in which border routers of autonomous systems (AS) adaptively deflect outbound traffic from a congested default path to an alternative path to significantly improve inter-domain traffic engineering (TE) and end-to-end throughput. With TED, every router only needs to examine the queue length of its own outgoing ports to orchestrate its deflection operation and ensure traffic forwarding is at line speed. It does not need to communicate or coordinate with other TED-capable routers or modify packet content, making TED incrementally deployable. Our evaluation shows that TED significantly increases the average throughput of traffic flows, and the improvement is comparable to directly upgrading router hardware and capacity. Finally, a prototype of TED on NetFPGA is also implemented.
Jun Li 0001, Ying Liu 0024, Dan Li 0001
IWQoS4
2014 Finding Constant from Change: Revisiting Network Performance Aware Optimizations on IaaS Clouds
abstract
Network performance aware optimizations have long been an effective approach to optimizing distributed applications on traditional network environments. However, the assumptions of network topology or direct use of several measurements of pair-wise network performance for optimizations are no longer valid on IaaS clouds. Virtualization hides network topology from users, and direct use of network performance measurements may not represent long-term performance. To enable existing network performance aware optimizations on IaaS clouds, we propose to decouple constant component from dynamic network performance while minimizing the difference by a mathematical method called RPCA (Robust Principal Component Analysis). We use the constant component to guide network performance aware optimizations and demonstrate the efficiency of our approach by adopting network aware optimizations for collective communications of MPI and generic topology mapping as well as two real-world applications, N-body and conjugate gradient (CG). Our experiments on Amazon EC2 and simulations demonstrate significant performance improvement on guiding the optimizations.
Yifan Gong 0003, Bingsheng He, Dan Li 0001
SC3
2014 LTTP: An LT-Code Based Transport Protocol for Many-to-One Communication in Data Centers
abstract
TCP has been widely adopted in current data centers to ensure reliable data delivery. However, recently TCP Incast was found to occur in many-to-one communications with barrier-synchronized requirement, where the TCP goodput drops dramatically. Previous solutions to TCP Incast either require updating the OS/hardware to support fine-grained timers, or smartly control utilization of the switch buffer to reduce the probability of buffer overflow and packet loss. In this paper we explore a different approach to support many-to-one communication in data center networks, which we call LTTP (LT-code based Transport Protocol). LTTP improves LT (Luby Transform) code to achieve reliable UDP-based transmission by exploiting data redundancy, and employs TFRC (TCP Friendly Rate Control) to adjust the traffic sending rates at servers. NS-2 based simulation shows that the goodput of LTTP never degrades with the increase of the number of servers in many-to-one communications, and LTTP significantly outperforms DCTCP when the number of servers is large. Simulation results also demonstrate that LTTP flows can fairly share bandwidth with TCP flows.
Changlin Jiang, Dan Li 0001, Mingwei Xu 0001
IEEE J. Sel. Areas Commun.2
2014 GreenDCN: A General Framework for Achieving Energy Efficiency in Data Center Networks
abstract
The popularization of cloud computing has raised concerns over the energy consumption that takes place in data centers. In addition to the energy consumed by servers, the energy consumed by large numbers of network devices emerges as a significant problem. Existing work on energy-efficient data center networking primarily focuses on traffic engineering, which is usually adapted from traditional networks. We propose a new framework to embrace the new opportunities brought by combining some special features of data centers with traffic engineering. Based on this framework, we characterize the problem of achieving energy efficiency with a time-aware model, and we prove its NP-hardness with a solution that has two steps. First, we solve the problem of assigning virtual machines (VM) to servers to reduce the amount of traffic and to generate favorable conditions for traffic engineering. The solution reached for this problem is based on three essential principles that we propose. Second, we reduce the number of active switches and balance traffic flows, depending on the relation between power consumption and routing, to achieve energy conservation. Experimental results confirm that, by using this framework, we can achieve up to 50 percent energy savings. We also provide a comprehensive discussion on the scalability and practicability of the framework.
Lin Wang 0015, Fa Zhang 0001, Jordi Arjona Aroca, Athanasios V. Vasilakos, Kai Zheng 0003, Chenying Hou, Dan Li 0001, Zhiyong Liu 0002
IEEE J. Sel. Areas Commun.7
2014 Reliable Multicast in Data Center Networks
abstract
Multicast benefits data center group communication in both saving network traffic and improving application throughput. Reliable packet delivery is required in data center multicast for data-intensive computations. However, existing reliable multicast solutions for the Internet are not suitable for the data center environment, especially with regard to keeping multicast throughput from degrading upon packet loss, which is norm instead of exception in data centers. We present RDCM, a novel reliable multicast protocol for data center network. The key idea of RDCM is to minimize the impact of packet loss on the multicast throughput, by leveraging the rich link resource in data centers. A multicast-tree-aware backup overlay is explicitly built on group members for peer-to-peer packet repair. The backup overlay is organized in such a way that it causes little individual repair burden, control overhead, as well as overall repair traffic. RDCM also realizes a window-based congestion control to adapt its sending rate to the traffic status in the network. Simulation results in typical data center networks show that RDCM can achieve higher application throughput and less traffic footprint than other representative reliable multicast protocols. We have implemented RDCM as a user-level library on Windows platform. The experiments on our test bed show that RDCM handles packet loss without obvious throughput degradation during high-speed data transmission, gracefully respond to link failure and receiver failure, and causes less than 10% CPU overhead to data center servers.
Dan Li 0001, Mingwei Xu 0001, Ying Liu 0024, Yong Cui 0001, Guihai Chen
IEEE Trans. Computers1
2014 Revisiting the Design of Mega Data Centers: Considering Heterogeneity Among Containers
abstract
In this paper, we revisit the design of mega data centers, which are usually built by a number of modularized containers. Due to technical innovation and vendor diversity, heterogeneity widely exists among data-center containers in practice. To embrace this issue, we propose uFix, which is a scalable, flexible, and modularized network architecture to interconnect heterogeneous data-center containers. The intercontainer connection rule in uFix is designed in such a way that it can flexibly scale to a huge number of servers with stable server/switch hardware settings. uFix allows modularized and fault-tolerant routing by completely decoupling intercontainer routing from intracontainer routing. We implement a software-based uFix prototype on a Linux platform. Both simulation and prototype-based experiment show that uFix enjoys high network capacity, gracefully handles server/switch failures, and causes lightweight CPU overhead onto data-center servers.
Dan Li 0001, Mingwei Xu 0001, Xiaoming Fu 0001
IEEE/ACM Trans. Netw.1
2013 PACE: Policy-Aware Application Cloud Embedding
abstract
The emergence of new capabilities such as virtualization and elastic (private or public) cloud computing infrastructures has made it possible to deploy multiple applications, on demand, on the same cloud infrastructure. A major challenge to achieve this possibility, however, is that modern applications are typically distributed, structured systems that include not only computational and storage entities, but also policy entities (e.g., load balancers, firewalls, intrusion prevention boxes). Deploying applications on a cloud infrastructure without the policy entities may introduce substantial policy violations and/or security holes. In this paper, we present PACE: the first systematic framework for Policy-Aware Application Cloud Embedding. We precisely define the policy-aware, cloud application embedding problem, study its complexity and introduce simple, efficient, online primal-dual algorithms to embed applications in cloud data centers. We conduct evaluations using data from a real, large campus network and a realistic data center topology to evaluate the feasibility and performance of PACE. We show that deployment in a cloud without considering in-network policies may lead to a large number of policy violations (e.g., using tree routing as a way to enforce in-network policies may observe up to 91% policy violations). We also show that our embedding algorithms are very efficient by comparing with a good online fractional embedding algorithm.
Li Erran Li, Vahid Liaghat, Mohammad Hajiaghayi, Dan Li 0001, Gordon T. Wilfong, Yang Richard Yang, Chuanxiong Guo
INFOCOM5
2013 Resource management in radio access and IP-based core networks for IMT Advanced and Beyond
Gang Su, Markus Hidell, Henrik Abrahamsson, Bengt Ahlgren, Dan Li 0001, Peter Sjödin, Voravit Tanyingyong, Ke Xu 0002
Sci. China Inf. Sci.5
2013 Expandable and Cost-Effective Network Structures for Data Centers Using Dual-Port Servers
abstract
A fundamental goal of data center networking is to efficiently interconnect a large number of servers with the low equipment cost. Several server-centric network structures for data centers have been proposed. They, however, are not truly expandable and suffer a low degree of regularity and symmetry. Inspired by the commodity servers in today's data centers that come with dual port, we consider how to build expandable and cost-effective structures without expensive high-end switches and additional hardware on servers except the two NIC ports. In this paper, two such network structures, called HCN and BCN, are designed, both of which are of server degree 2. We also develop the low overhead and robust routing mechanisms for HCN and BCN. Although the server degree is only 2, HCN can be expanded very easily to encompass hundreds of thousands servers with the low diameter and high bisection width. Additionally, HCN offers a high degree of regularity, scalability, and symmetry, which conform to the modular designs of data centers. BCN is the largest known network structure for data centers with the server degree 2 and network diameter 7. Furthermore, BCN has many attractive features, including the low diameter, high bisection width, large number of node-disjoint paths for the one-to-one traffic, and good fault-tolerant ability. Mathematical analysis and comprehensive simulations show that HCN and BCN possess excellent topological properties and are viable network structures for data centers.
Deke Guo, Tao Chen 0013, Dan Li 0001, Mo Li 0001, Yunhao Liu 0001, Guihai Chen
IEEE Trans. Computers3
2013 Dynamic Scheduling for Wireless Data Center Networks
abstract
Unbalanced traffic demands of different data center applications are an important issue in designing data center networks (DCN). In this paper, we present our exploratory investigation on a hybrid DCN solution of utilizing wireless transmissions in DCNs. Our work aims to solve the congestion problem caused by a few hot nodes to improve the global performance. We model the wireless transmissions in DCN by considering both the wireless interference and the adaptive transmission rate. Besides, both throughput and job completion time are considered to measure the impact of wireless transmissions on the global performance. Based on the model, we formulate the problem of channel allocation as an optimization problem. We also design an approximation algorithm with an approximation bound of 1/2 and a genetic algorithm to address the scheduling problem. A series of simulations are performed to evaluate and demonstrate the effectiveness of our wireless DCN scheme.
Yong Cui 0001, Hongyi Wang 0004, Xiuzhen Cheng, Dan Li 0001, Antti Ylä-Jääski
IEEE Trans. Parallel Distributed Syst.4
2013 IP-Geolocation Mapping for Moderately Connected Internet Regions
abstract
Most IP-geolocation mapping schemes [14], [16], [17], [18] take delay-measurement approach, based on the assumption of a strong correlation between networking delay and geographical distance between the targeted client and the landmarks. In this paper, however, we investigate a large region of moderately connected Internet and find the delay-distance correlation is weak. But we discover a more probable rule - with high probability the shortest delay comes from the closest distance. Based on this closest-shortest rule, we develop a simple and novel IP-geolocation mapping scheme for moderately connected Internet regions, called GeoGet. In GeoGet, we take a large number of webservers as passive landmarks and map a targeted client to the geolocation of the landmark that has the shortest delay. We further use JavaScript at targeted clients to generate HTTP/Get probing for delay measurement. To control the measurement cost, we adopt a multistep probing method to refine the geolocation of a targeted client, finally to city level. The evaluation results show that when probing about 100 landmarks, GeoGet correctly maps 35.4 percent clients to city level, which outperforms current schemes such as GeoLim [16] and GeoPing [14] by 270 and 239 percent, respectively, and the median error distance in GeoGet is around 120 km, outperforming GeoLim and GeoPing by 37 and 70 percent, respectively.
Dan Li 0001, Chuanxiong Guo, Yunxin Liu 0001, Zhi-Li Zhang, Yongguang Zhang
IEEE Trans. Parallel Distributed Syst.1
2012 SIONA: A service and information oriented network architecture
abstract
The Internet is a great hit in human history. However, it has evolved greatly from its original incarnation. Content distribution is playing a central role in today's Internet, which makes it difficult for the conventional host-to-host communication to meet the ever-increasing demands. In this paper, we present a novel “service and information oriented network architecture” (SIONA). The key aspect of SIONA is the name-based two-dimensional routing paradigm that provides scalable routing, caching and content delivery. We argue that SIONA solves the problems of mobile Internet by naturally supporting mobility, and provides network layer P2P for massive data distribution. Evaluation is conducted to investigate its caching and mobility performance.
Zhongxing Ming, Mingwei Xu 0001, Chunmei Xia, Dan Li 0001, Dan Wang 0002
ICC4
2012 Towards bandwidth guarantee in multi-tenancy cloud computing networks
abstract
To efficiently utilize their infrastructure and thus increase their revenue, cloud providers need mechanisms to provide resource allocation and performance isolation for different tenants in the shared platform. In particular, network bandwidth sharing is a critical yet still an open problem to most cloud providers. In this paper, we study the problem of virtual machine (VM) allocation under the consideration of providing bandwidth guarantees. We first propose an online allocation algorithm for tenants with homogeneous bandwidth demand, which improves on the accuracy of existing algorithms. Subsequently, we extend it to handle heterogeneous bandwidth demand. Extensive simulations show that our algorithm makes much more efficient utilization of the network resource than existing algorithms, and performs close to the optimal offline allocation.
Jing Zhu 0007, Dan Li 0001, Hongnan Liu
ICNP2
2012 ESM: Efficient and Scalable Data Center Multicast Routing
abstract
Multicast benefits group communications in saving network traffic and improving application throughput, both of which are important for data center applications. However, the technical trend of data center design poses new challenges for efficient and scalable multicast routing. First, the densely connected networks make traditional receiver-driven multicast routing protocols inefficient in multicast tree formation. Second, it is quite difficult for the low-end switches widely used in data centers to hold the routing entries of massive multicast groups. In this paper, we propose ESM, an efficient and scalable multicast routing scheme for data center networks. ESM addresses the challenges above by exploiting the feature of modern data center networks. Based on the regular topology of data centers, ESM uses a source-to-receiver expansion approach to build efficient multicast trees, excluding many unnecessary intermediate switches used in receiver-driven multicast routing. For scalable multicast routing, ESM combines both in-packet Bloom Filters and in-switch entries to make the tradeoff between the number of multicast groups supported and the additional bandwidth overhead. Simulations show that ESM saves 40%$\sim$50% network traffic and doubles the application throughputs compared to receiver-driven multicast routing, and the combination routing scheme significantly reduces the number of in-switch entries required. We implement ESM on a Linux platform. The experimental results further demonstrate that ESM can well support online tree building for large-scale groups with churns, and the overhead of the combination forwarding engine is light-weighted.
Dan Li 0001, Yuanjie Li, Sen Su, Jiangwei Yu
IEEE/ACM Trans. Netw.1
2011 Building mega data center from heterogeneous containers
abstract
Data center containers are regarded as the basic units to build mega data centers. In practice, heterogeneity exists among data center containers, because of technical innovation and vendor diversity. In this paper, we propose uFix, a scalable, flexible and modularized network architecture to interconnect heterogeneous data center containers. The inter-container connection rule in uFix is designed in such a way that it can flexibly scale to a huge number of servers with stable server/switch hardware settings. uFix allows modularized and fault-tolerant routing by completely decoupling inter-container routing from intra-container routing. We implement a software-based uFix stack on the Linux platform. Simulation and experiment results show that uFix enjoys high network capacity, gracefully handles server/switch failures, and brings light-weight CPU overhead onto data center servers.
Dan Li 0001, Mingwei Xu 0001, Xiaoming Fu 0001
ICNP1
2011 BCN: Expansible network structures for data centers using hierarchical compound graphs
abstract
A fundamental challenge in data centers is how to design networking structures for efficiently interconnecting a large number of servers. Several server-centric structures have been proposed, but are not truly expansible and suffer low degree of regularity and symmetry. To address this issue, we propose two novel structures called HCN and BCN, which utilize hierarchical compound graphs to interconnect large population of servers each with two ports only. They own two topological advantages, i.e., the expansibility and equal degree. In addition, HCN offers high degree of regularity, scalability and symmetry, which well conform to the modular design of data centers. Moreover, a BCN of level one in each dimension involves more servers than FiConn with server degree 2 and diameter 7, and is large enough for a single data center. Mathematical analysis and comprehensive simulations show that BCN possesses excellent topology properties and is a viable network structure for data centers.
Deke Guo, Tao Chen 0013, Dan Li 0001, Yunhao Liu 0001, Xue (Steve) Liu, Guihai Chen
INFOCOM3
2011 RDCM: Reliable data center multicast
abstract
Multicast benefits data center group communication in both saving network traffic and improving application throughput. The SLA (Service Level Agreement) of cloud service requires the computation correctness of distributed applications, translating to the requirement of reliable Multicast delivery. In this paper we present RDCM, a novel reliable Multicast approach for data center network. The key idea of RDCM is to minimize the impact of packet loss on the Multicast performance, by leveraging the rich link resource in data centers. A Multicast-tree-aware backup overlay is purposely built on group members for peer-to-peer packet repair. Riding on Unicast, packet repair not only achieves complete repair isolation, but also has high probability to bypass the pathological links in the Multicast tree where packet loss occurs. The backup overlay is organized in such a way that it causes little individual repair burden, control overhead, as well as overall repair traffic. We have implemented RDCM as a user-level library on Windows platform. The experiments on our test bed show that RDCM handles packet loss without obvious throughput degradation during high-speed data transmission.
Dan Li 0001, Mingwei Xu 0001, Ming-Chen Zhao, Chuanxiong Guo, Yongguang Zhang, Min-You Wu
INFOCOM1
2011 Exploring efficient and scalable multicast routing in future data center networks
abstract
Multicast benefits group communications in saving network traffic and improving application throughput, both of which are important for data center applications. However, the technical trend of future data center design poses new challenges for efficient and scalable Multicast routing. First, the densely connected networks make traditional receiver-driven Multicast routing protocols inefficient in Multicast tree formation. Second, it is quite difficult for the low-end switches largely used in data centers to hold the routing entries of massive Multicast groups.
Dan Li 0001, Jiangwei Yu, Junbiao Yu
INFOCOM1
2011 Impact of user selfishness in construction action on the streaming quality of overlay multicast
Dan Li 0001, Yong Cui 0001, Jiangchuan Liu, Ke Xu 0002
Comput. Networks1
2011 Scalable and cost-effective interconnection of data-center servers using dual server ports
abstract
The goal of data-center networking is to interconnect a large number of server machines with low equipment cost while providing high network capacity and high bisection width. It is well understood that the current practice where servers are connected by a tree hierarchy of network switches cannot meet these requirements. In this paper, we explore a new server-interconnection structure. We observe that the commodity server machines used in today's data centers usually come with two built-in Ethernet ports, one for network connection and the other left for backup purposes. We believe that if both ports are actively used in network connections, we can build a scalable, cost-effective interconnection structure without either the expensive higher-level large switches or any additional hardware on servers. We design such a networking structure called FiConn. Although the server node degree is only 2 in this structure, we have proven that FiConn is highly scalable to encompass hundreds of thousands of servers with low diameter and high bisection width. We have developed a low-overhead traffic-aware routing mechanism to improve effective link utilization based on dynamic traffic state. We have also proposed how to incrementally deploy FiConn.
Dan Li 0001, Chuanxiong Guo, Kun Tan 0001, Yongguang Zhang, Songwu Lu
IEEE/ACM Trans. Netw.1
2011 Defending Against Distance Cheating in Link-Weighted Application-Layer Multicast
abstract
Application-layer multicast (ALM) has recently emerged as a promising solution for diverse group-oriented applications. Unlike dedicated routers in IP multicast, the autonomous end-hosts are generally unreliable and even selfish. A strategic host might cheat about its private information to affect protocol execution and, in turn, to improve its individual benefit. Specifically, in a link-weighted ALM protocol where the hosts measure the distances from their neighbors and accordingly construct the ALM topology, a selfish end-host can easily intercept the measurement message and exaggerate the distances to other nodes, so as to reduce the probability of being a relay. Such distance cheating, rarely happening in IP multicast, can significantly impact the efficiency and stability of the ALM topology. To defend against this kind of cheating, we present a Vickrey–Clarke–Groves (VCG)-based cheat-proof mechanism in this paper. We demonstrate a practical mapping from the utility, payment, and welfare of a VCG mechanism to the link-weighted ALM context. Based on this, we further discuss practical issues for implementing the cheat-proof mechanism—specifically, a trustworthy distributed algorithm for payment computation. Performance analyses show that the overheads of the computation, storage, and communication of our implementation are controlled at low levels, and extensive simulations further testify the implementation's effectiveness. Although there are other similar studies in this area, the contribution of our cheat-proof mechanism and its implementation primarily lies in two aspects. On one hand, we first explicitly solve the distance cheating problem in link-weighted ALM since its proposal by mapping the VCG mechanism to link-weighted ALM context. On the other hand, our distributed implementation can not only effectively defend against distance cheating, but can also avoid the potential cheating behaviors when selfish ALM nodes fulfill the cheat-proof mechanism itself.
Dan Li 0001, Jiangchuan Liu, Yong Cui 0001, Ke Xu 0002
IEEE/ACM Trans. Netw.1
2010 WIND: A scalable and lightweight network topology service for peer-to-peer applications
abstract
We present an Internet-scale network topology information (NTI) service named WIND for localizing P2P traffic. Central to WIND are the two simple ideas: 1) obtaining NTI directly from routing infrastructures, and 2) leveraging existing, widely deployed DNS caches for NTI delivery. WIND fulfills the fidelity, flexibility and scalability requirement of an effective NTI service. WIND is deployed in CERNET. We conduct extensive trace-driven emulations on PlanetLab. Experimental results confirm the effectiveness of the WIND service.
Hongqiang Liu, Yongqiang Xiong, CongXiao Bao, Xing Li 0001, Guobin Shen, Dan Li 0001
NOMS6
2009 MDCube: a high performance network structure for modular data center interconnection
abstract
Shipping-container-based data centers have been introduced as building blocks for constructing mega-data centers. However, it is a challenge on how to interconnect those containers together with reasonable cost and cabling complexity, due to the fact that a mega-data center can have hundreds or even thousands of containers and the aggregate bandwidth among containers can easily reach tera-bit per second. As a new inner-container server-centric network architecture, BCube [9] interconnects thousands of servers inside a container and provides high bandwidth support for typical traffic patterns. It naturally serves as a building block for mega-data center.
Guohan Lu, Dan Li 0001, Chuanxiong Guo, Yongguang Zhang
CoNEXT3
2009 FiConn: Using Backup Port for Server Interconnection in Data Centers
abstract
The goal of data center networking is to interconnect a large number of server machines with low equipment cost, high and balanced network capacity, and robustness to link/server faults. It is well understood that, the current practice where servers are connected by a tree hierarchy of network switches cannot meet these requirements (Fares et al., 2008 and Guo et al., 2008). In this paper, we explore a new server-interconnection structure. We observe that the commodity server machines used in today's data centers usually come with two built-in Ethernet ports, one for network connection and the other left for backup purpose. We believe that, if both ports are actively used in network connections, we can build a low-cost interconnection structure without the expensive higher-level large switches. Our new network design, called FiConn, utilizes both ports and only the low-end commodity switches to form a scalable and highly effective structure. Although the server node degree is only two in this structure, we have proven that FiConn is highly scalable to encompass hundreds of thousands of servers with low diameter and high bisection width. The routing mechanism in FiConn balances different levels of links. We have further developed a low-overhead traffic-aware routing mechanism to improve effective link utilization based on dynamic traffic state. Simulation results have demonstrated that the routing mechanisms indeed achieve high networking throughput.
Dan Li 0001, Chuanxiong Guo, Kun Tan 0001, Songwu Lu
INFOCOM1
2009 BCube: a high performance, server-centric network architecture for modular data centers
abstract
This paper presents BCube, a new network architecture specifically designed for shipping-container based, modular data centers. At the core of the BCube architecture is its server-centric network structure, where servers with multiple network ports connect to multiple layers of COTS (commodity off-the-shelf) mini-switches. Servers act as not only end hosts, but also relay nodes for each other. BCube supports various bandwidth-intensive applications by speeding-up one-to-one, one-to-several, and one-to-all traffic patterns, and by providing high network capacity for all-to-all traffic.
Chuanxiong Guo, Guohan Lu, Dan Li 0001, Yunfeng Shi, Chen Tian 0001, Yongguang Zhang, Songwu Lu
SIGCOMM3
2009 Defending Against Buffer Map Cheating in DONet-Like P2P Streaming
abstract
Data-driven overlay network (DONet)-like P2P system is especially suitable to support live stream applications, since its data structure can tolerate node dynamics quite well. However, optimal streaming demands the cooperation of individual nodes. If selfish nodes cheat about their buffer maps to reduce the forwarding burden, the overall streaming quality would be negatively affected. To defend against this kind of cheating, we design a trustworthy service-differentiation based incentive mechanism with low complexity in this paper. The mechanism is composed of the service-differentiation algorithm and the contribution-evaluation algorithm. Compared with other studies in this area, the primary characteristic of our mechanism lies in two aspects. Firstly, the contribution of each node is evaluated considering the features of live streaming, not just by the transferring bytes. Secondly, the potential cheating behavior of overlay nodes during the fulfillment of incentive algorithms can be avoided, which is usually not considered by other similar studies. Extensive simulations suggest that the algorithms are indeed effective for defending against buffer map cheating in DONet-like P2P streaming.
Dan Li 0001, Yong Cui 0001
IEEE Trans. Multim.1
2007 Truthful Streaming in Selfish DONet
abstract
Data-driven overlay network (DONet) is especially suitable for live stream because it can tolerant node dynamics well. However, optimal streaming demands the cooperation of individual nodes. If selfish nodes in DONet cheat about their buffer maps to reduce the forwarding burden, the overall streaming quality might be negatively affected. To defend this kind of cheating behavior, we design a trustworthy service- differentiation based incentive mechanism with low complexity in this paper. The mechanism is composed of the service- differentiation algorithm and the contribution-evaluation algorithm. Compared with other studies in this area, the primary characteristic of our mechanism lies in two aspects. Firstly, the contribution of each node is evaluated considering the characteristic of live stream, not just by the transferring bytes. Secondly, the potential cheating behavior of overlay nodes during the fulfillment of incentive algorithms can be defended, which is usually not considered by other studies.
Dan Li 0001, Yong Cui 0001
ICC1
2007 QoS-Aware Streaming in Overlay Multicast Considering the Selfishness in Construction Action
abstract
Most existing overlay multicast proposals have assumed that the nodes are cooperative and thus focus on the global topology optimization. However, a unique and important characteristic of overlay nodes is that, as application-layer agents, they can be selfish with their own interests. To achieve better quality-of-service (QoS) or to minimize forwarding overhead, an overlay node can behave selfishly in the information collection or in the overlay construction. While the former has recently been investigated, the impact of selfishness in the construction action remains unclear. In this paper, we present the first systematic study on the impact of selfishness in both tree and mesh overlay construction. Our investigation considers multiple QoS measures for streaming applications, including stream latency, resolution, and continuity. Our contribution is twofold: first, we analyze how for selfish overlay nodes to choose a construction-action policy to optimize their individual multi-metric QoS. Second, we demonstrate that the selfishness-aware policy for the construction action is consistent with the QoS optimization for the global multicast session, but not vice versa. The implication is significant: A globally optimal overlay construction itself can be vulnerable to individual selfishness; but, following our directions, we can design an overlay that is both globally optimal and selfish-resistant.
Dan Li 0001, Yong Cui 0001, Jiangchuan Liu
INFOCOM1
2006 Segment-sending Schedule in Data-driven Overlay Network
abstract
Data-driven overlay network is suitable for live-event streaming, because it can provide relatively-continuous streaming even in dynamic environment. In terms of improving streaming quality, prior work covered membership management, buffer map exchange, segment requesting schedule, etc. In this paper, we address the problem of segment-sending schedule on the segment-providing node, which may also affect the streaming quality. The schedule methods we discuss include FIFO schedule, lower-sequence favored schedule, and higher-sequence favored schedule. Simulation results show that if users care playing continuity much more than playing delay, the higher-sequence favored schedule brings the best streaming quality; however, if users care playing delay much more than playing continuity, lower-sequence favored schedule is the preferred choice. Through this work, we find another way to improve streaming quality in data-driven overlay network.
Dan Li 0001, Yong Cui 0001, Ke Xu 0002
ICC1
2006 Stability of ALM Tree with selfish receivers: A simulation study
Dan Li 0001, Yong Cui 0001
Comput. Commun.2
2005 Impact of receiver cheating on the stability of ALM tree
abstract
Application layer multicast (ALM) is an effective supplement to IP multicast, but it has the potential trouble of trust on end systems. For instance, multicast receivers may cheat in order to obtain a better position in the multicast tree. Receiver cheating may transform the multicast tree, and lead to its instability. We establish the cheating model of ALM receivers and analyze the stability of ALM tree when receiver cheating occurs. Simulation results show that receiver cheating has considerably negative effects on the stability of ALM tree. This discovery brings forward an issue in ALM study, that is, we should take receiver cheating into consideration to maintain a stable ALM tree when designing ALM protocols.
Dan Li 0001, Yong Cui 0001, Ke Xu 0002
GLOBECOM1