EDBT 2026 Demo / reviewers in the wild / expert
Dan Li 0001
dblp:48/4185-1
· DBLP profile ↗
130ranked-venue papers
22as first author
59since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 92 · 16 first-author · 46 since 2021Systems, architecture and hardware · 18 · 4 first-author · 3 since 2021Security and privacy · 7 · 7 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OptiFlow: Towards LLM-Driven Optimization of Collective Communication Algorithms
Ziyue Yang 0002, Kaihui Gao, Shuai Wang 0028, Li Chen 0008, Zhixiong Niu, Ran Shu 0001, Wenxue Cheng, Peng Cheng 0005, Yongqiang Xiong, Dan Li 0001 |
APNet | 11 |
| 2026 | CATS: Predictive-Feedback Adaptive Load Balancing for Computing-Aware Traffic Steering
Yuxiang Shang, Tao Sun 0010, Dan Li 0001, Zhenping Hu, Lu Lu 0016, Chengjiang Wen, Yantao Han, Li Chen 0008, Huijuan Yao, Peng Liu 0047 |
ICC | 3 |
| 2026 | Accurate and Stable AS Relationship Inference via Trusted Seeds and Semi-Supervised Learning
Siyuan Teng, Lancheng Qin, Li Chen 0008, Dan Li 0001 |
INFOCOM | 4 |
| 2026 | A Large-Scale IPv6-Based Measurement of the Starlink NetworkabstractLow Earth Orbit (LEO) satellite networks have attracted considerable attention for their ability to deliver global, low-latency broadband Internet services. In this paper, we present a large-scale measurement study of the Starlink network, the largest LEO satellite constellation to date. We first propose an efficient method for discovering active Starlink user routers, identifying approximately 5.98 million IPv6 addresses across 208 regions in 165 countries. Compared to general-purpose IPv6 target generation algorithms, our router-centric approach achieves near-complete coverage and, to the best of our knowledge, yields the most comprehensive known set of active IPv6 addresses for Starlink user routers. Based on the discovered user routers, we further propose an efficient method for mapping the Starlink backbone network and uncover a topology consisting of 49 Points of Presence (PoPs) interconnected by 98 links. We conduct a detailed statistical analysis of active Starlink user routers and PoPs, and further characterize the IPv6 address assignment strategy adopted by the Starlink network. Finally, we analyze the latency of Starlink user routers, propose a method to distinguish different types of users within the same region using outside-in measurement, and identify the ongoing V2 Mini satellite deployment as a potential driver of the performance improvements. The dataset of the Starlink backbone network is publicly available at https://ki3.org.cn/#/starlink-network. Bingsen Wang, Shuai Wang 0028, Li Chen 0008, Jinwei Zhao, Dan Li 0001, Yong Jiang 0001 |
INFOCOM | 6 |
| 2026 | HeraClass: Towards Open-World Network Flow Classification via Traffic-Language Mapping
Ni Jin, Libin Liu 0001, Yukai Miao, Li Chen 0008, Dan Li 0001, Xizheng Wang, Xiuting Xu, Baojiang Cui |
IWQoS | 5 |
| 2026 | OSAVRoute: Advancing Outbound Source Address Validation Deployment Detection with Non-Cooperative Measurement
Shuai Wang 0028, Li Chen 0008, Dan Li 0001, Lancheng Qin |
NDSS | 4 |
| 2026 | BayWatch: Practical Internet-Scale Topology Monitoring with Dynamic Bayesian Estimation
Zhongxu Guan, Shuai Wang 0028, Li Chen 0008, Zhaoteng Yan, Jiaye Lin, Dan Li 0001, Yong Jiang 0001, Yingxin Wang |
NSDI | 6 |
| 2026 | CCEval: Accurately and Confidently Evaluating Performance Metrics of Congestion Control Algorithms for Datacenter Networks
Tianfeng Liu, Kaihui Gao, Li Chen 0008, Dan Li 0001, Jin Guang, Vincent Liu 0001, Yiwei Zhang 0016, Ni Jin |
NSDI | 4 |
| 2026 | Supercharging Packet-level Network Simulation of Large Model Training via Memoization and Fast-Forwarding
Kaihui Gao, Li Chen 0008, Dan Li 0001, Yiwei Zhang 0016, Fei Gui, Yitao Xing, Wenjia Wei, Bingyang Liu |
NSDI | 4 |
| 2026 | PReCCL: Performant and Resilient Collective Communication via Integrated Inband Telemetry and Workload ReallocationabstractModern collective communication libraries (CCLs) execute a collective communication task (CCT) by decomposing it into multiple sub-tasks, each mapped to a specific Virtual Topology (VT), which is an ordered graph of GPUs (e.g., a ring or a tree), to maximize parallelism and link utilization. As AI training scales to larger clusters, network anomalies (congestion and failures) are unavoidable, and a single straggling VT can delay the entire CCT. Existing solutions either rely on low-level transport-layer solutions which lacks a cross-sub-task perspective, or static CCL scheduling, failing to adapt to the dynamic and heterogeneous networks. Kaihui Gao, Li Chen 0008, Fei Gui, Dan Li 0001, Jiamin Cao |
SIGCOMM | 7 |
| 2026 | Networked Agent Memory and Causality Representation: Experiences towards Interpretable Cloud-Scale Root-Causing
Yanyu Ren, Xianshang Lin, Chenxu Wang 0007, Li Chen 0008, Shuai Wang 0028, Kaihui Gao, Dan Li 0001, Chen Tian 0001, Yunguang Li, Ennan Zhai |
SIGCOMM | 7 |
| 2026 | Open the Floodgates in a Digital Twin: Experiences of Building Spillway for 100M+-User Signaling Storms in Cellular Core NetworkabstractSignaling storms threaten cellular core networks when synchronized reconnection attempts from massive numbers of devices trigger cascading, metastable overloads. Existing defenses rely on manual, static configurations of local overload controls, which ignore serial dependencies among heterogeneous network elements. We present Spillway, a digital-twin-driven system that automates global signaling-flood mitigation. Spillway introduces a hierarchical defense architecture that enforces altruistic throttling, allowing upstream nodes to shed load before downstream bottlenecks collapse. To evaluate candidate configurations, Spillway uses CN-DES, a domain-specific discrete-event simulator with a vectorized kernel. By aggregating users that share protocol states, CN-DES decouples simulation cost from user count and simulates regional-scale storms involving tens of millions of users in minutes, achieving a 60× speedup over traditional simulation while preserving fidelity. Spillway then uses heteroscedastic evolutionary Bayesian optimization to search a large, non-convex parameter space. We report on a five-year deployment in the world's largest 5G Standalone network. During real incidents, including application anomalies and RAN failures, networks using Spillway-optimized configurations experienced substantially fewer user fallbacks than predicted under legacy configurations; post-incident analysis confirms that pre-deployed parameters kept all network elements within safe operating bounds. Hongtao Xie 0006, Jianmin Liu, Li Chen 0008, Dan Li 0001, Mineng Fu, Xi Chen 0026 |
SIGCOMM | 6 |
| 2026 | I2BGP: A Privacy-Preserving Intra-AS State-Assisted Inter-AS Routing SchemeabstractBGP is the most widely employed inter-AS routing protocol, connecting millions of ASes worldwide. While it is possible to select the egress for outgoing flows based on administrators’ configurations, such schemes are localized due to the privacy of the intra-AS network state. TheAS_Pathfield of BGP records all crossed ASes, which can be used to prevent routing loops and select paths,i.e., selecting the minimum AS-hop path among available paths. Although this scheme is simple, effective, and offers a certain degree of global perspective, selecting paths at AS granularity ignores the transmission performance within each intra-AS, which may result in selecting non-optimal routing paths. To enable the use of private intra-AS data for inter-AS routing, we proposed a privacy-preserving intra-AS state-assisted inter-AS routing scheme, which can select optimal inter-AS paths without disclosing specific intra-AS state data. Specifically, we added an additional BGP header field to carry path performance features and designed a three-step data masking mechanism to protect intra-AS state data, enabling the selection of inter-AS paths with intra-AS state awareness. I2BGP has been deployed in the Greater Bay Area Future Network and a large-scale network simulator based on real network topologies. The results show that I2BGP outperforms BGP in terms of specified forwarding hops, delay, and bandwidth metrics. Peizhuang Cong, Yuchao Zhang 0004, Jun Wang 0178, Wendong Wang 0003, Tong Yang 0003, Dan Li 0001, Ke Xu 0002 |
IEEE Trans. Netw. | 6 |
| 2026 | Example Generalizing Network Configuration Synthesizer via Graph-Informed Large Language Models
Jianmin Liu, Li Chen 0008, Dan Li 0001, Yukai Miao, Liyu Ma |
IEEE Trans. Netw. | 3 |
| 2025 | A Theoretical Framework for Quantitative Evaluation of Padding Defenses against Website Fingerprinting
Dan Li 0001 |
APNet | 3 |
| 2025 | Towards Automatic Network Diagram ComprehensionabstractNetwork Diagram Comprehension (NDC) is a vital task for networking professionals, offering essential insights into network topology and configurations. However, NDC remains a labor-intensive process heavily reliant on human expertise, with existing tools falling short in addressing this challenge. It is critical to develop an Automatic NDC (ANDC) system that ensures high faithfulness and completeness in information extraction while supporting practical, end-to-end NDC applications. Moreover, a comprehensive dataset and benchmark are necessary to systematically evaluate and drive the progress of ANDC.In this work, we introduce Layered Extractor of Network Diagrams (LEND), the first ANDC system designed to comprehensively and faithfully extract and utilize information from network diagrams. LEND employs a three-stage pipeline: (1) a layer extractor to decompose diagrams and identify key elements with a denoising cascade, (2) an inter-layer combiner to reconstruct entity relations with positional and domain knowledge, and (3) a task-specific interpreter for networking applications.To support this effort, we develop two extensive NDC datasets comprising over 4,000 network diagrams and icons from diverse sources, along with the first benchmark to evaluate ANDC systems across three distinct metrics. Empirical experiments demonstrate that LEND outperforms existing methods by achieving at 1.21– 5.10× better faithfulness and completeness, and improves its capability as a NetOps engineer by 30.5% on the Cisco Certified Network Associate (CCNA) exam. Yanyu Ren, Yukai Miao, Li Chen 0008, Dan Li 0001, Xizheng Wang, Yu Bai 0021 |
ICNP | 4 |
| 2025 | Measuring the Time Source Vulnerabilities in the NTP EcosystemabstractPrecise timekeeping is crucial for the dependable functioning and security of multiple Internet infrastructures, such as TLS certificates. Although the Network Time Protocol (NTP) is widely used for time synchronization across devices, it has several security vulnerabilities. Network Time Security (NTS) offers server authentication and integrity verification to protect against man-in-the-middle attacks. However, NTS does not address issues related to erroneous time sources. Zhentian Huang, Shuai Wang 0028, Li Chen 0008, Dan Li 0001 |
IMC | 4 |
| 2025 | SAIP: Accurate Detection of Anycast Servers with the Rise of Regional Anycast
Shuai Wang 0028, Li Chen 0008, Dan Li 0001 |
INFOCOM | 4 |
| 2025 | Transcending Cost-Quality Tradeoff in Agent Serving via Session-AwarenessabstractLarge Language Model (LLM) agents are capable of task execution across various domains by autonomously interacting with environments and refining LLM responses based on feedback.
However, existing model serving systems are not optimized for the unique demands of serving agents. Compared to classic model serving, agent serving has different characteristics:
predictable request pattern, increasing quality requirement, and unique prompt formatting. We identify a key problem for agent serving: LLM serving systems lack session-awareness. They neither perform effective KV cache management nor precisely select the cheapest yet competent model in each round.
This leads to a cost-quality tradeoff, and we identify an opportunity to surpass it in an agent serving system.
To this end, we introduce AgServe for AGile AGent SERVing.
AgServe features a session-aware server that boosts KV cache reuse via Estimated-Time-of-Arrival-based eviction and in-place positional embedding calibration, a quality-aware client that performs session-aware model cascading through real-time quality assessment, and a dynamic resource scheduler that maximizes GPU utilization.
With AgServe, we allow agents to select and upgrade models during the session lifetime, and to achieve similar quality at much lower costs, effectively transcending the tradeoff. Extensive experiments on real testbeds demonstrate that AgServe (1) achieves comparable response quality to GPT-4o at a 16.5\% cost. (2) delivers 1.8$\times$ improvement in quality relative to the tradeoff curve. Yanyu Ren, Li Chen 0008, Dan Li 0001, Xizheng Wang, Yukai Miao, Yu Bai 0021 |
NeurIPS | 3 |
| 2025 | Accelerating Design Space Exploration for LLM Training Systems with Multi-experiment Parallel Simulation
Fei Gui, Kaihui Gao, Li Chen 0008, Dan Li 0001, Vincent Liu 0001, Hongbing Yang, Dian Xiong |
NSDI | 4 |
| 2025 | CEGS: Configuration Example Generalizing Synthesizer
Jianmin Liu, Li Chen 0008, Dan Li 0001, Yukai Miao |
NSDI | 3 |
| 2025 | Resolving Packets from Counters: Enabling Multi-scale Network Traffic Super Resolution via Composable Large Traffic Model
Xizheng Wang, Libin Liu 0001, Li Chen 0008, Dan Li 0001, Yukai Miao, Yu Bai 0021 |
NSDI | 4 |
| 2025 | SimAI: Unifying Architecture Design and Performance Tuning for Large-Scale Large Language Model Training with Scalability and Precision
Xizheng Wang, Qingxu Li, Yichi Xu, Dan Li 0001, Li Chen 0008, Heyang Zhou, Linkang Zheng, Yikai Zhu, Yang Liu 0245, Kun Qian 0021, Kunling He, Ennan Zhai, Dennis Cai, Binzhang Fu |
NSDI | 5 |
| 2025 | Discovering Millions of New Nodes and Links in the Internet by Challenging the Uniformity Assumption in Multipath DetectionabstractMultipath Detection Algorithms (MDAs) are proposed to discover Internet topology in the presence of load balancing (LB). Existing methods assume uniformity in the load-balancing responses (LBR), i.e., responses from the successors of a LB router. However, we reveal that only 20% of the cases exhibit uniformity in the Internet. This finding significantly challenges the completeness of the Internet topology discovered using current MDAs. In this paper, we propose a novel system BayMuDA, that can estimate LBR distributions and calculate the minimum number of probes needed to statistically discover all nodes and links within a given hop. The validation on controlled topologies shows that BayMuDA discovers at least 85%/73% of nodes/links in ~90% of the cases. Our Internet-wide measurement results indicate that BayMuDA can discover millions of Internet nodes and links obscured by the state-of-the-art MDA algorithm, D-Miner, due to uneven responses. Zhongxu Guan, Shuai Wang 0028, Li Chen 0008, Zhaoteng Yan, Jiaye Lin, Dan Li 0001, Yong Jiang 0001, Yingxin Wang |
SIGCOMM | 6 |
| 2025 | From ATOP to ZCube: Automated Topology Optimization Pipeline and A Highly Cost-Effective Network Topology for Large Model TrainingabstractThe development of large language models (LLMs) poses new challenges in data center network topology design. To assist in exploring topology design, we propose ATOP, an Automated Topology Optimization Pipeline, which models network topology as a set of hyperparameters, enabling the discovery of potential topologies. With various optimization algorithms and customizable optimization objectives, ATOP achieves automated topology optimization on a scale of tens of thousands of GPUs. We apply ATOP on network topologies for 256, 1024, 4096, and 16384 GPUs, optimizing performance under LLMs training traffic patterns, collective communication performance, fault tolerance, and network cost. We also evaluate ATOP in different scenarios: building, optimizing, and expanding a data center. From ATOP's results, we discover a new topology — ZCube, which reaches the highest cost-effectiveness across various GPU scales. Simulation results show that ZCube, compared to the previous state-of-the-art topologies, including Rail-optimized Fat-tree (ROFT), Rail-only, and HPN, improves end-to-end LLM training speed by 3% to 7% and reduces network hardware costs by 26% to 46%. We also construct ZCube on a real-world testbed. Results show that ZCube reduces hardware costs by 25% compared to Rail-Optimized Topology while maintaining the same all-reduce and all-to-all performance. Dan Li 0001, Li Chen 0008, Dian Xiong, Kaihui Gao, Yiwei Zhang 0016, Menglei Zhang, Bochun Zhang, Zhuo Jiang, Jianxi Ye, Haibin Lin |
SIGCOMM | 2 |
| 2025 | The Digital Cybersecurity Expert: How Far Have We Come?abstractThe increasing deployment of large language models (LLMs) in the cybersecurity domain underscores the need for effective model selection and evaluation. However, traditional evaluation methods often overlook specific cybersecurity knowledge gaps that contribute to performance limitations. To address this, we develop CSEBenchmark, a fine-grained cyber-security evaluation framework based on 345 knowledge points expected of cybersecurity experts. Drawing from cognitive science, these points are categorized into factual, conceptual, and procedural types, enabling the design of 11,050 tailored multiple-choice questions. We evaluate 12 popular LLMs on CSEBenchmark and find that even the best-performing model achieves only 85.42% overall accuracy, with particular knowledge gaps in the use of specialized tools and uncommon commands. Different LLMs have unique knowledge gaps. Even large models from the same family may perform poorly on knowledge points where smaller models excel. By identifying and addressing specific knowledge gaps in each LLM, we achieve up to an 84% improvement in correcting previously incorrect predictions across three existing benchmarks for two cybersecurity tasks. Furthermore, our assessment of each LLM's knowledge alignment with specific cybersecurity roles reveals that different models align better with different roles, such as GPT-4o for the Google Senior Intelligence Analyst and Deepseek-V3 for the Amazon Privacy Engineer. These findings underscore the importance of aligning LLM selection with the specific knowledge requirements of different cybersecurity roles for optimal performance. Dawei Wang 0021, Geng Zhou, Yu Bai 0021, Li Chen 0008, Ting Qin, Dan Li 0001 |
SP | 8 |
| 2025 | Your Shield is My Sword: A Persistent Denial-of-Service Attack via the Reuse of Unvalidated Caches in DNSSEC Validation
Shuai Wang 0028, Li Chen 0008, Dan Li 0001 |
USENIX Security Symposium | 4 |
| 2024 | ProphetFuzz: Fully Automated Prediction and Fuzzing of High-Risk Option Combinations with Only Documentation via Large Language ModelabstractVulnerabilities related to option combinations pose a significant challenge in software security testing due to their vast search space. Previous research primarily addressed this challenge through mutation or filtering techniques, which inefficiently treated all option combinations as having equal potential for vulnerabilities, thus wasting considerable time on non-vulnerable targets and resulting in low testing efficiency. In this paper, we utilize carefully designed prompt engineering to drive the large language model (LLM) to predict high-risk option combinations (i.e., more likely to contain vulnerabilities) and perform fuzz testing automatically without human intervention. We developed a tool called ProphetFuzz and evaluated it on a dataset comprising 52 programs collected from three related studies. The entire experiment consumed 10.44 CPU years. ProphetFuzz successfully predicted 1748 high-risk option combinations at an average cost of only \8.69 per program. Results show that after 72 hours of fuzzing, ProphetFuzz discovered 364 unique vulnerabilities associated with 12.30% of the predicted high-risk option combinations, which was 32.85% higher than that found by state-of-the-art in the same timeframe. Additionally, using ProphetFuzz, we conducted persistent fuzzing on the latest versions of these programs, uncovering 140 vulnerabilities, with 93 confirmed by developers and 21 awarded CVE numbers. Dawei Wang 0021, Geng Zhou, Li Chen 0008, Dan Li 0001, Yukai Miao |
CCS | 4 |
| 2024 | Robust or Risky: Measurement and Analysis of Domain Resolution DependencyabstractDNS relies on domain delegation for good scalability, where domains delegate their resolution service to authoritative nameservers. However, such delegations lead to complex inter-dependencies between DNS zones. While a complex dependency might improve the robustness of domain resolution, it could also introduce security risks unexpectedly. In this work, we perform a large-scale measurement on nearly 217M domains to analyze their resolution dependencies at both zone level and infrastructure level. According to our analysis, domains under country-code TLDs and new generic TLDs generally present more complicated dependency relationships. For robustness consideration, popular domains prefer to configure more complex dependencies. However, the centralization of nameserver hosting and the silent outsourcing of DNS providers could lead to severe false redundancy at infrastructure level. Worse, considerable domain configurations in the wild are "not robust but risky": a more complex dependency may also indicate more vulnerabilities, e.g., domains with a 2× higher dependency complexity have a 2.87× larger proportion suffering from the hijacking risk brought by lame delegation. Shuai Wang 0028, Dan Li 0001 |
INFOCOM | 3 |
| 2024 | Understanding Route Origin Validation (ROV) Deployment in the Real World and Why MANRS Action 1 Is Not Followed
Lancheng Qin, Li Chen 0008, Dan Li 0001, Honglin Ye |
NDSS | 3 |
| 2024 | dRR: A Decentralized, Scalable, and Auditable Architecture for RPKI Repository
Dan Li 0001, Li Chen 0008, Qi Li 0002, Sitong Ling |
NDSS | 2 |
| 2024 | RedTE: Mitigating Subsecond Traffic Bursts with Real-time and Distributed Traffic EngineeringabstractInternet traffic bursts usually happen within a second, thus conventional burst mitigation methods ignore the potential of Traffic Engineering (TE). However, our experiments indicate that a TE system, with a sub-second control loop latency, can effectively alleviate burst-induced congestion. TE-based methods can leverage network-wide tunnel-level information to make globally informed decisions (e.g., balancing traffic bursts among multiple paths). Our insight in reducing control loop latency is to let each router make local TE decisions, but this introduces the key challenge of minimizing performance loss compared to centralized TE systems. Fei Gui, Dan Li 0001, Li Chen 0008, Kaihui Gao, Congcong Min, Yi Wang 0004 |
SIGCOMM | 3 |
| 2024 | Deep Distributional Reinforcement Learning-Based Adaptive Routing With Guaranteed Delay BoundsabstractReal-time applications that require timely data delivery over wireless multi-hop networks within specified deadlines are growing increasingly. Effective routing protocols that can guarantee real-time QoS are crucial, yet challenging, due to the unpredictable variations in end-to-end delay caused by unreliable wireless channels. In such conditions, the upper bound on the end-to-end delay, i.e., worst-case end-to-end delay, should be guaranteed within the deadline. However, existing routing protocols with guaranteed delay bounds cannot strictly guarantee real-time QoS because they assume that the worst-case end-to-end delay is known and ignore the impact of routing policies on the worst-case end-to-end delay determination. In this paper, we relax this assumption and propose DDRL-ARGB, an Adaptive Routing with Guaranteed delay Bounds using Deep Distributional Reinforcement Learning (DDRL). DDRL-ARGB adopts DDRL to jointly determine the worst-case end-to-end delay and learn routing policies. To accurately determine worst-case end-to-end delay, DDRL-ARGB employs a quantile regression deep Q-network to learn the end-to-end delay cumulative distribution. To guarantee real-time QoS, DDRL-ARGB optimizes routing decisions under the constraint of worst-case end-to-end delay within the deadline. To improve traffic congestion, DDRL-ARGB considers the network congestion status when making routing decisions. Extensive results show that DDRL-ARGB can accurately calculate worst-case end-to-end delay, and can strictly guarantee real-time QoS under a small tolerant violation probability against two state-of-the-art routing protocols. Jianmin Liu, Dan Li 0001, Yongjun Xu 0001 |
IEEE/ACM Trans. Netw. | 2 |
| 2024 | Decentralized and Incentivized Federated Learning: A Blockchain-Enabled Framework Utilising Compressed Soft-Labels and Peer ConsistencyabstractFederated Learning (FL) has emerged as a powerful paradigm in Artificial Intelligence, facilitating the parallel training of Artificial Neural Networks on edge devices while safeguarding data privacy. Nonetheless, to encourage widespread adoption, Federated Learning Frameworks (FLFs) must tackle (i) the power imbalance between a central authority and its participants, and (ii) the challenge of equitably measuring and incentivizing contributions. Existing approaches to decentralize and incentivize FL processes are hindered by (i) computational overhead and (ii) uncertainty in contribution assessment [1]), limiting FL's scalability beyond use cases where trust between participants and the server is established. This work introduces a cutting-edge, blockchain-enabled federated learning framework that incorporates Federated Knowledge Distillation (FD) with compressed 1-bit soft-labels, aggregated through a smart contract. Furthermore, we present the Peer Truth Serum for Federated Distillation (PTSFD), which cultivates an incentive-compatible ecosystem by rewarding honest participation based on an implicit yet effective comparison of worker contributions. The primary innovation stems from its lightweight architecture that simultaneously promotes decentralization and incentivization, addressing critical challenges in contemporary FL approaches. Leon Witt, Usama Zafar, KuoYeh Shen, Felix Sattler, Dan Li 0001, Wojciech Samek |
IEEE Trans. Serv. Comput. | 5 |
| 2023 | sRDMA: A General and Low-Overhead Scheduler for RDMAabstractRemote Direct Memory Access (RDMA) has been widely deployed in data centers to improve application performance. However, the characteristic of RDMA to deliver messages in order cannot meet the emerging requirements of applications for scheduling messages within an RDMA connection, making RDMA unable to be fully utilized. Some works try to schedule the data to be transferred in specific applications before delivering to RDMA, or distribute messages to different connections. However, these approaches tightly couple scheduling logic with application logic and may result in high scheduling overhead. Xizheng Wang, Shuai Wang 0028, Dan Li 0001 |
APNet | 3 |
| 2023 | Impact of International Submarine Cable on Internet Routing
Honglin Ye, Shuai Wang 0028, Dan Li 0001 |
INFOCOM | 3 |
| 2023 | BGL: GPU-Efficient GNN Training by Optimizing Graph Data I/O and Preprocessing
Tianfeng Liu, Yangrui Chen, Dan Li 0001, Chuan Wu 0001, Yibo Zhu 0001, Yanghua Peng, Hongzheng Chen, Chuanxiong Guo |
NSDI | 3 |
| 2023 | Demo: NetVision: Efficient Visualization Front-End for Packet-level Discrete-Event Network SimulationabstractVisualization of network simulation is an essential tool for network practitioners. However, the front-end of existing network simulators often fails to deliver satisfactory performance when dealing with modern network scales and interface speed. In this paper, we propose NetVision, an efficient visualization front-end of network simulation based on the Unity engine, which is commonly used for video game and virtual reality development. NetVision offers flow-level visualization of network behavior and performances. Then, through parallel optimization, NetVision supports real-time visualization for large-scale high-speed networks. Kaihui Gao, Li Chen 0008, Dan Li 0001, Vincent Liu 0001, Xizheng Wang, Lu Lu 0016 |
SIGCOMM | 3 |
| 2023 | DONS: Fast and Affordable Discrete Event Network Simulation with Automatic ParallelizationabstractDiscrete Event Simulation (DES) is an essential tool for network practitioners. Unfortunately, existing DES simulators cannot achieve satisfactory performance at the scale of modern networks. Recent work has attempted to address these challenges by reducing the traffic processed via novel approximation techniques; however, we argue in this paper that much of the slowdown of existing DES simulators is due to their underlying software architecture. Kaihui Gao, Li Chen 0008, Dan Li 0001, Vincent Liu 0001, Xizheng Wang, Lu Lu 0016 |
SIGCOMM | 3 |
| 2023 | Poster: Q-Scanner: A Fast Scanning Tool for Large-Scale SSL/TLS Configurations MeasurementabstractSecure Sockets Layer (SSL) and Transport Layer Security (TLS) protocols are used to encrypt data, protect privacy, and authenticate. However, the security of SSL/TLS itself depends on its configurations. While some scanning tools are used to measure SSL/TLS configurations, their performance is far from meeting the requirement of large-scale measurements. In this paper, we propose a fast SSL/TLS configuration scanning tool, Q-Scanner, which can generate a lightweight scanning solution based on the characteristics of the configurations to be scanned. The experiment shows Q-Scanner achieves a speedup of over 30,000 times compared to SSL Pulse without loss of accuracy. Shuai Wang 0028, Dan Li 0001 |
SIGCOMM | 3 |
| 2023 | Light: A Compatible, high-performance and scalable user-level network stack
Dan Li 0001, Huiyou Jiang, Du Lin, Jinkun Geng, K. K. Ramakrishnan, Kai Zheng 0003 |
Comput. Networks | 2 |
| 2023 | DIT and Beyond: Interdomain Routing With Intradomain Awareness for IIoTabstractAlong with the ever-increasing amount of data generated from industrial devices, the cross domain [also known as autonomous systems (ASs)] data transmission problem has attracted more and more attention in the Industrial Internet of Things (IIoT). As mature and widely used interdomain routing protocols, border gateway protocol-based solutions often take the number of domains (i.e., AS hops) of each path as a criterion to make routing decisions, which is simple and effective. However, such protocols can only meet the reachability requirements while ignoring the performance requirements. That is, the path with the minimum AS hops will be selected to carry flows, even if the actual performance of this path does not meet the transmission requirements due to the unawareness of intradomain information on that path. But it is not impractical to directly access intradomain information for making better routing decisions given data privacy concerns. In this article, we propose M-DIT, which can make interdomain routing decisions with the assistance of desensitized intradomain information for multiple-requirement transmissions. To do so, we design a homomorphic encrypted-based private number comparison scheme to export intradomain information securely and, thus, assist in routing decisions. The results of some experiments based on five real topologies (ATMnet,Claranet,Compuserve,NSFnet, andPeer1) with thousands of interdomain flows demonstrate that M-DIT reduced flow completion time by about 60% or selected high bandwidth paths flexibly for interdomain routing for IIoT scenarios. Peizhuang Cong, Yuchao Zhang 0004, Wendong Wang 0003, Xiangyang Gong, Tong Yang 0003, Dan Li 0001, Ke Xu 0002 |
IEEE Internet Things J. | 7 |
| 2023 | Decentral and Incentivized Federated Learning Frameworks: A Systematic Literature ReviewabstractThe advent of federated learning (FL) has sparked a new paradigm of parallel and confidential decentralized machine learning (ML) with the potential of utilizing the computational power of a vast number of Internet of Things (IoT), mobile, and edge devices without data leaving the respective device, thus ensuring privacy by design. Yet, simple FL frameworks (FLFs) naively assume an honest central server and altruistic client participation. In order to scale this new paradigm beyond small groups of already entrusted entities toward mass adoption, FLFs must be: 1) truly decentralized and 2) incentivized to participants. This systematic literature review is the first to analyze FLFs that holistically apply both, the blockchain technology to decentralize the process and reward mechanisms to incentivize participation. 422 publications were retrieved by querying 12 major scientific databases. After a systematic filtering process, 40 articles remained for an in-depth examination following our five research questions. To ensure the correctness of our findings, we verified the examination results with the respective authors. Although having the potential to direct the future of distributed and secure artificial intelligence, none of the analyzed FLFs is production ready. The approaches vary heavily in terms of use cases, system design, solved issues, and thoroughness. We provide a systematic approach to classify and quantify differences between FLFs, expose limitations of current works and derive future directions for research in this novel domain. Leon Witt, Mathis Heyer, Kentaroh Toyoda, Wojciech Samek, Dan Li 0001 |
IEEE Internet Things J. | 5 |
| 2023 | Buffer-Based High-Coverage and Low-Overhead Request Event Monitoring in the CloudabstractRequest latency directly affects the performance of modern cloud applications. Due to various causes in hosts and networks, requests can suffer from request latency anomalies (RLAs), which may violate the Service-Level Agreement. However, existing performance monitoring tools have incomplete coverage and inconsistent semantics for monitoring requests and cannot accurately diagnose RLAs. This paper presentsBufScope, a high-coverage and low-overhead request event monitoring system, which monitorsbuffersto capture most RLA-related abnormal events with consistent request-level semantics in the end-to-end datapath of request. First,BufScopemodels the datapath of request as a buffer chain and defines events based on three properties of buffers, so as toend-to-end monitorthe root causes of RLA. Then, to achieveconsistent semanticsfor captured events,BufScopedesigns a request-level semantics injection mechanism to make events captured in networks have the victim requests’ ID. Finally,BufScopeoffloads the semantics operations and event collection in software to SmartNICs forlow CPU overhead. We have implementedBufScopeon commodity SmartNICs and programmable switches. Evaluation results show thatBufScopecan diagnose 98% RLAs with < 0.08% network bandwidth overhead and 0.6% application throughput decline. Kaihui Gao, Chen Sun 0005, Shuai Wang 0028, Dan Li 0001, Yu Zhou 0008, Hongqiang Harry Liu, Lingjun Zhu, Ming Zhang 0005, Lu Lu 0016 |
IEEE/ACM Trans. Netw. | 4 |
| 2023 | Dependable Virtualized Fabric on Programmable Data PlaneabstractIn modern multi-tenant data centers, each tenant desires reassuring dependability from the virtualized network fabric – bandwidth guarantee with work conservation, bounded tail latency and resilient reachability. However, the slow convergence of prior works under network dynamics and uncertainties can hardly provide the dependability for tenants. Further, state-of-the-art load balance schemes are guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works. In this paper, we propose vFab, a dependable virtualized fabric framework which can (1) quickly detect network failure in data plane, (2) explicitly select proper paths for all flows, and (3) converge to ideal bandwidth allocation at sub-millisecond. The core idea of vFab is to leverage the programmable data plane to build a fusion of an active edge (e.g., NIC) and an informative core (e.g., switch), where the core sends link status and tenant information to the edge via telemetry to help the latter make a timely and accurate decision on path selection and traffic admission. We fully implement vFab with commodity SmartNICs and programmable switches. Extensive evaluations show that vFab can keep bandwidth guarantee with high bandwidth utilization, low and bounded latency, and resilient reachability under various network scenarios with limited overhead. Application-level experiments show that vFab can improve QPS by$2.4\times $and cut tail latency by$10\times $compared to the alternatives. Kaihui Gao, Shuai Wang 0028, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Tao Sun 0010 |
IEEE/ACM Trans. Netw. | 4 |
| 2022 | Bandwidth-efficient Microburst Measurement in Large-scale Datacenter NetworksabstractMicroburst measurement is essential for diagnosing and mitigating performance problems in datacenter networks. The key is to efficiently identify the flows that contribute the most to queue buildup. However, because the existing microburst measurement systems capture packet-level information, they incur significant bandwidth overhead. We present BurstScope, a bandwidth-efficient microburst measurement system that can profile the microburst characteristics and the contributing flows. BurstScope detects the microburst-involved packets in egress pipeline, then aggregates the measurement granularity from packet level to flow level by an invertible sketch. Finally, by carefully partitioning the measurement and statistic tasks between the data and control plane, we generate only one telemetry packet for each microburst. We have implemented BurstScope on Barefoot Tofino switches. Testbed-based evaluations show that BurstScope keeps low bandwidth overhead (< 0.02%) and high identification accuracy (> 97%). Compared with the state-of-the-art system, BurstScope can reduce 60 × bandwidth overhead. Kaihui Gao, Dan Li 0001, Shuai Wang 0028 |
APNet | 2 |
| 2022 | EndGraph: An Efficient Distributed Graph Preprocessing SystemabstractGraph processing mainly includes two stages, namely, preprocessing and algorithm execution. Most previous proposals for performance enhancement of graph processing systems focus on the algorithm execution stage, and simple ignore the preprocessing overhead. However, in this work, we argue that the cost of preprocessing can not be ignored since the preprocessing time is much longer than the algorithm execution time in state-of-the-art systems.We propose EndGraph, a distributed graph preprocessing system, to improve preprocessing performance. Firstly, for graph partitioning, we find existing systems either assign imbalanced preprocessing workloads or spend too much time on graph partitioning. Hence, EndGraph proposes a novel chunk-based partition algorithm to balance preprocessing workloads and achieve theoretical lower bound of time complexity. Secondly, for graph construction (converting data layout from edge array to adjacency list), existing systems use counting sort, which is not efficient for computation and communication. EndGraph employs a novel two-level graph construction method by carefully decoupling the graph construction into intra-machine and inter-machine construction. Our extensive evaluation results show that, compared with five state-of-the-art systems, LFGraph, PowerLyra, PowerGraph, D-Galois, and Gemini, EndGraph can improve the preprocessing performance up to 35.76 ×(from 4.72×). To show the generality of EndGraph, we integrate it with D-Galois and Gemini, and it improves the end-to-end (including preprocessing and algorithm execution) graph processing performance up to 7.44× (from 2.96×). Tianfeng Liu, Dan Li 0001 |
ICDCS | 2 |
| 2022 | Break the Blackbox! Desensitize Intra-domain Information for Inter-domain RoutingabstractAlong with the ever-increasing amount of data generated from edge networks, cross domain (also known as Autonomous Systems, AS) transmission problem has attracted more and more attention. As mature and widely used inter-domain routing protocols, BGP-based solutions often use the number of domains (i.e. AS hops) of each path to make inter-domain routing decisions, which is simple and effective, but usually can not get the optimal routing results due to the lack of real state/information within ASes. These protocols choose the path with less AS hops as the forwarding path, even if the total latency or cost of the domains on this path is higher. While to solve this problem, directly access to intra-domain information as the assistance to make routing decisions is impractical due to data privacy.In this paper, we propose DIT, which makes near-optimal inter-domain routing decisions with desensitized intra-domain information. To do so, we design a homomorphic encrypted-based private number comparison scheme to export intra-domain information securely and thus assist in routing decisions. We conduct a series of experiments according to five real network topologies with nearly 900 simulated flows, and the results show that DIT reduces the number of forwarding hops by about 45% in average and reduces flow completion time by about 60%. Peizhuang Cong, Yuchao Zhang 0004, Wendong Wang 0003, Xiangyang Gong, Tong Yang 0003, Dan Li 0001, Ke Xu 0002 |
IWQoS | 8 |
| 2022 | Buffer-based End-to-end Request Event Monitoring in the Cloud
Kaihui Gao, Chen Sun 0005, Shuai Wang 0028, Dan Li 0001, Yu Zhou 0008, Hongqiang Harry Liu, Lingjun Zhu, Ming Zhang 0005 |
NSDI | 4 |
| 2022 | Elixir: A High-performance and Low-cost Approach to Managing Hardware/Software Hybrid Flow Tables Considering Flow Burstiness
Yanshu Wang, Dan Li 0001, Yuanwei Lu |
NSDI | 2 |
| 2022 | Predictable vFabric on informative data planeabstractIn multi-tenant data centers, each tenant desires reassuring predictability from the virtual network fabric - bandwidth guarantee, work conservation, and bounded tail latency. Achieving these goals simultaneously relies on rapid and precise traffic admission. However, the slow convergence (tens of milliseconds) of prior works can hardly satisfy the increasingly rigorous performance demand under dynamic traffic patterns. Further, state-of-the-art load balance schemes are all guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works. Shuai Wang 0028, Kaihui Gao, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Ming Zhang 0005 |
SIGCOMM | 4 |
| 2022 | Themis: Accelerating the Detection of Route Origin Hijacking by Distinguishing Legitimate and Illegitimate MOAS
Lancheng Qin, Dan Li 0001 |
USENIX Security Symposium | 2 |
| 2022 | Impact of Synchronization Topology on DML Performance: Both Logical Topology and Physical TopologyabstractTo tackle the increasingly larger training data and models, researchers and engineers resort to multiple servers in a data center for distributed machine learning (DML). On one hand, DML enables us to leverage the computation power of multiple servers, which can effectively accelerate those computation-intensive tasks. On the other hand, DML also incurs significant communication cost due to parameter synchronization among these servers. In this paper, we want to explore the impact of synchronization topology, including both logical topology and physical topology, on the DML performance. First, we revisit the existing logical topologies, e.g., parameter server and ring allreduce, for parameter synchronization, and we find that theseflatsynchronization topologies is inefficient when running a large-scale DML training. Therefore, we propose a hierarchical parameter synchronization topology, called HiPS, which can achieve efficient parameter synchronization even on a large scale. Then, we compare two representative physical network topologies, namely, Fat-Tree and BCube. Based on our analyses, BCube has many advantages over Fat-Tree, e.g., higher bandwidth, better load balance, and lower hardware cost. The simulation results also show that BCube is more friendly to RDMA. Relying on the advantages of HiPS and BCube, the GST of “HiPS+BCube” is 12% ~ 70% lower than other combinations. Moreover, when the cluster size increases from 16 to 1024, the performance of “HiPS+BCube” only drops by 6.5%, while the performance of “Ring+BCube” drops by 44.6%. Hence, we believe “HiPS+BCube” is the optimal solution to benefit DML in large scale. Shuai Wang 0028, Jinkun Geng, Dan Li 0001 |
IEEE/ACM Trans. Netw. | 3 |
| 2021 | Mu: An Efficient, Fair and Responsive Serverless Framework for Resource-Constrained Edge CloudsabstractServerless computing platforms simplify development, deployment, and automated management of modular software functions. However, existing serverless platforms typically assume an over-provisioned cloud, making them a poor fit for Edge Computing environments where resources are scarce. In this paper we propose a redesigned serverless platform that comprehensively tackles the key challenges for serverless functions in a resource constrained Edge Cloud. Viyom Mittal, Shixiong Qi, Ratnadeep Bhattacharya, Xiaosu Lyu, Sameer G. Kulkarni, Dan Li 0001, Jinho Hwang, K. K. Ramakrishnan, Timothy Wood 0001 |
SoCC | 7 |
| 2021 | Fast and Robust Online Traffic Classification Supporting Unseen ApplicationsabstractOnline traffic classification is a fundamental toolkit in network management, such as QoS and network security. The speed and generalization ability of online classification are two requirements that need to be satisfied simultaneously. However, existing methods may suffer from generalization degradation on the traffic with unseen applications which are constantly emerging in the network, due to the feature distribution drift (FDD) caused by their non-robust feature engineering approaches. Based on Deep Metric Learning which can restrict the distances between samples explicitly and clustering algorithm which can learn multiple clusters within each category, this paper presents Robot, a fast and robust online traffic classification system. At its core, Robot leverages two building blocks to classify high-speed traffic flows: 1) For fast classification, Fast model classifies traffic and detects FDD samples simultaneously based on only one packet. 2) For robust classification, once FDD samples are detected, a flow collector will be triggered to collect flows and then Robust model, a multi-center model, will further identify them based on the hybrid of packet-level and flow-level features. Our comprehensive experiments demonstrate that Robot can achieve comparative classification speed and better generalization ability on the mixed traffic datasets with seen and unseen applications (with the FDD detection accuracy of up to 85.8% and with nearly 10% improvement in classification accuracy), compared with the state-of-the-art methods. Dan Li 0001, Kaihui Gao |
GLOBECOM | 2 |
| 2021 | FastKeeper: A Fast Algorithm for Identifying Top-k Real-time Large FlowsabstractPrecise identification of large flows is a critical task in network traffic measurement. Previous works focus on identification of elephant flows (i.e. large flows from the beginning of the measurement). However, we generally observe that the flow rates change periodically and abruptly. In addition, the large flows may become small flows over time. Thus, elephant flows are not equal to the real-time large flows, and previous works cannot be used for the identification of the real-time large flows that is more meaningful for modern network applications. Nevertheless, identification of real-time large flows is challenging in that it requires accurate measurement of real-time flow rates and timely replacement of flows that have become small in the measurement data structure. In this paper, we propose FastKeeper to identify real-time large flows with a primary goal of simultaneously achieving low over-head, high performance and high accuracy. FastKeeper employs a sliding-window-based algorithm for accurate measurement of real-time flow rates and a bitmap-voting algorithm for timely replacement of flows that have become small in the measurement data structure. We evaluate it on DPDK using the traces from an operator network, and the evaluation demonstrates that it achieves high accuracy (98%) and processing throughput (25.33Mpps). Yanshu Wang, Dan Li 0001 |
GLOBECOM | 2 |
| 2021 | Analyzing Open-Source Serverless Platforms: Characteristics and Performance (S)abstractServerless computing is increasingly popular because of its lower cost and easier deployment.Several cloud service providers (CSPs) offer serverless computing on their public clouds, but it may bring the vendor lock-in risk.To avoid this limitation, many open-source serverless platforms come out to allow developers to freely deploy and manage functions on self-hosted clouds.However, building effective functions requires much expertise and thorough comprehension of platform frameworks and features that affect performance.It is a challenge for a service developer to differentiate and select the appropriate serverless platform for different demands and scenarios.Thus, we elaborate the frameworks and event processing models of four popular open-source serverless platforms and identify their salient idiosyncrasies.We analyze the root causes of performance differences between different service exporting and auto-scaling modes on those platforms.Further, we provide several insights for future work, such as auto-scaling and metric collection.Index Terms-cloud computing, Sameer G. Kulkarni, K. K. Ramakrishnan, Dan Li 0001 |
SEKE | 4 |
| 2021 | Sphinx: A transport protocol for high-speed and lossy mobile networks
Dan Li 0001, Wenfei Wu, K. K. Ramakrishnan, Jinkun Geng, Fanzhao Wang, Kai Zheng 0003 |
Comput. Networks | 2 |
| 2021 | Accelerating End-to-End Deep Learning Workflow With Codesign of Data Preprocessing and SchedulingabstractIn this article, we investigate the performance bottleneck of existing deep learning (DL) systems and propose DLBooster to improve the running efficiency of deploying DL applications on GPU clusters. At its core, DLBooster leverages two-level optimizations to boost the end-to-end DL workflow. On the one hand, DLBooster selectively offloads some key decoding workloads to FPGAs to provide high-performance online data preprocessing services to the computing engine. On the other hand, DLBooster reorganizes the computational workloads of training neural networks with the backpropagation algorithm and schedules them according to their dependencies to improve the utilization of GPUs at runtime. Based on our experiments, we demonstrate that compared with baselines, DLBooster can improve the image processing throughput by 1.4× - 2.5× and reduce the processing latency by 1/3 in several real-world DL applications and datasets. Moreover, DLBooster consumes less than 1 CPU core to manage FPGA devices at runtime, which is at least 90 percent less than the baselines in some cases. DLBooster shows its potential to accelerate DL workflows in the cloud. Dan Li 0001, Binyao Jiang, Jinkun Geng, Wei Bai 0001, Yongqiang Xiong |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | CEFS: compute-efficient flow scheduling for iterative synchronous applicationsabstractIterative Synchronous Applications (ISApps) are popular in today's data centers, represented by distributed deep learning (DL) training. In ISApps, multiple nodes carry out the computing task iteratively, with globally synchronizing the results in each iteration. To increase the scaling efficiency of ISApps, in this paper we propose a new flow scheduling approach, called CEFS. CEFS saves the waiting time of computing nodes from two aspects. For a single node, flows with data which can trigger earlier computation at the node are assigned with higher priority; among nodes, flows towards slower nodes are assigned with higher priority. Shuai Wang 0028, Dan Li 0001, Jiansong Zhang 0001, Wei Lin 0016 |
CoNEXT | 2 |
| 2020 | Fela: Incorporating Flexible Parallelism and Elastic Tuning to Accelerate Large-Scale DMLabstractDistributed machine learning (DML) has become the common practice in industry, because of the explosive volume of training data and the growing complexity of training model. Traditional DML follows data parallelism but causes significant communication cost, due to the huge amount of parameter transmission. The recently emerging model-parallel solutions can reduce the communication workload, but leads to load imbalance and serious straggler problems. More importantly, the existing solutions, either data-parallel or model-parallel, ignore the nature of flexible parallelism for most DML tasks, thus failing to fully exploit the GPU computation power. Targeting at these existing drawbacks, we propose Fela, which incorporates both flexible parallelism and elastic tuning mechanism to accelerate DML. In order to fully leverage GPU power and reduce communication cost, Fela adopts hybrid parallelism and uses flexible parallel degrees to train different parts of the model. Meanwhile, Fela designs token-based scheduling policy to elastically tune the workload among different workers, thus mitigating the straggler effect and achieve better load balance. Our comparative experiments show that Fela can significantly improve the training throughput and outperforms the three main baselines (i.e. dataparallel, model-parallel, and hybrid-parallel) by up to 3.23×, 12.22×, and 1.85× respectively. Jinkun Geng, Dan Li 0001, Shuai Wang 0028 |
ICDE | 2 |
| 2020 | Geryon: Accelerating Distributed CNN Training by Network-Level Flow SchedulingabstractIncreasingly rich data sets and complicated models make distributed machine learning more and more important. However, the cost of extensive and frequent parameter synchronizations can easily diminish the benefits of distributed training across multiple machines. In this paper, we present Geryon, a network-level flow scheduling scheme to accelerate distributed Convolutional Neural Network (CNN) training. Geryon leverages multiple flows with different priorities to transfer parameters of different urgency levels, which can naturally coordinate multiple parameter servers and prioritize the urgent parameter transfers in the entire network fabric. Geryon requires no modification in CNN models and does not affect the training accuracy. Based on the experimental results of four representative CNN models on a testbed of 8 GPU servers, Geryon achieves up to 95.7% scaling efficiency even with 10GbE bandwidth. In contrast, for most models, the scaling efficiency of vanilla TensorFlow is no more than 37% and that of TensorFlow with parameter partition and slicing is around 80%. In terms of training throughput, Geryon enhanced with parameter partition and slicing achieves up to 4.37x speedup, where the flow scheduling algorithm itself achieves up to 1.2x speedup over parameter partition and slicing. Shuai Wang 0028, Dan Li 0001, Jinkun Geng |
INFOCOM | 2 |
| 2020 | Incorporating Intra-flow Dependencies and Inter-flow Correlations for Traffic Matrix PredictionabstractTraffic matrix (TM) prediction is essential for effective traffic engineering and network management. Based on our analysis of real traffic traces from Wide Area Network, the traffic flows in TM are both time-varying (i.e. with intra-flow dependencies) and correlated with each other (i.e. with inter-flow correlations). However, most existing works in TM prediction ignore inter-flow correlations. In this paper, we propose a novel Attention-based Convolutional Recurrent Neural Network (ACRNN) model to capture both intra-flow dependencies and inter-flow correlations. ACRNN mainly contains two components: 1) Correlational Modeling employs attention-based convolutional structures to capture the correlation of any two flows in TMs; 2) Temporal Modeling uses attention-based recurrent structures to model the long-term temporal dependencies of each flow, and then predicts TMs according inter-flow correlations and intra-flow dependencies. Experiments on two real-world datasets show that, when predicting the next TM, ACRNN model reduces the Mean Squared Error by up to 44.8% and reduces the Mean Absolute Error by up to 30.6%, compared to state-of-the-art method; and the gap is even larger when predicting the next multiple TMs. Besides, simulation results demonstrate that ACRNN's accurate prediction can help traffic engineering to mitigate traffic congestion. Kaihui Gao, Dan Li 0001, Li Chen 0008, Jinkun Geng, Fei Gui |
IWQoS | 2 |
| 2020 | Managing Multicast Membership for Software Defined Data Center NetworkabstractIn this paper we design DCMA, a novel multicast membership management scheme for data center networks. Unlike traditional protocols like IGMP/MLD, DCMA leverages the characteristics of data center multicast application and the emerging software defined networking (SDN) technique to manage multicast members in an easier and better way. Multicast application master delivers the membership to the SDN controller, and the SDN controller assigns the group addresses in a coordinated way to minimize the forwarding table size in switches. In particular, by formulating the multicast group address allocation problem and capturing the relationship among forwarding entries in different switches, we design both a batch algorithm and an incremental algorithm to allocate the multicast group addresses. Evaluations based on real-world traces show that, DCMA can save almost half multicast forwarding entries in switches compared with random multicast address allocation. Dan Li 0001, Jing Zhu 0007, Hongnan Liu, Kai Chen 0005 |
VTC Fall | 2 |
| 2020 | A Scalable, High-Performance, and Fault-Tolerant Network Architecture for Distributed Machine LearningabstractIn large-scale distributed machine learning (DML), the network performance between machines significantly impacts the speed of iterative training. In this paper we propose BML, a scalable, high-performance and fault-tolerant DML network architecture on top of Ethernet and commodity devices. BML builds on BCube topology, and runs a fully-distributed gradient synchronization algorithm. Compared to a Fat-Tree network with the same size, a BML network is expected to take much less time for gradient synchronization, for both low theoretical synchronization time and its benefit to RDMA transport. With server/link failures, the performance of BML degrades in a graceful way. Experiments of MNIST and VGG-19 benchmarks on a testbed with 9 dual-GPU servers show that, BML reduces the job completion time of DML training by up to 56.4% compared with Fat-Tree running state-of-the-art gradient synchronization algorithm. Dan Li 0001, Jinkun Geng, Yanshu Wang, Shuai Wang 0028, Shutao Xia |
IEEE/ACM Trans. Netw. | 2 |
| 2019 | Accelerating Distributed Machine Learning by Smart Parameter ServerabstractParameter Server (PS)-based architecture is widely applied in distributed machine learning (DML), but it is still an open issue how to improve the DML performance in this frame-work. Existing works mainly focus on the view of workers. In this paper, we tackle this problem from another perspective, by leveraging the central control on the PS. Specifically, we propose SmartPS, which transforms the passive role of PS in traditional DML and fully exploits the intelligence of PS. Firstly, the PS holds the global view of parameter dependency, facilitating it to update workers' parameters selectively and proactively. Secondly, the PS records the workers' speeds, and prioritizes parameter transmission to narrow the gap between stragglers and fast workers. Thirdly, the PS considers the parameter dependency in consecutive training iterations, and opportunistically blocks unnecessary pushes from workers. We conduct comparative experiments with two typical benchmarks, Matrix Factorization (MF) and PageRank (PR). The experimental results prove that, compared with all the baseline algorithms (i.e. standard BSP, ASP and SSP), SmartPS can reduce the overall training time by 65.7%~84.9%, with the same training accuracy. Jinkun Geng, Dan Li 0001, Shuai Wang 0028 |
APNet | 2 |
| 2019 | Rima: An RDMA-Accelerated Model-Parallelized Solution to Large-Scale Matrix FactorizationabstractMatrix factorization (MF) is a fundamental technique in machine learning and data mining, which gains wide application in many fields. When the matrix becomes large, MF cannot be processed on a single machine. Considering this, many distributed SGD algorithms (e.g. DSGD) have been developed to solve large-scale MF on multiple machines in a model-parallel way. Existing distributed algorithms are primarily implemented under Map/Reduce or PS (parameter server)-based architectures, which incur significant communication overheads. Besides, existing solutions cannot well embrace the benefit of RDMA/RoCE transport and suffer from scalability problems. Targeting at these drawbacks, we propose Rima, which uses ring-based model parallelism to solve large-scale MF with higher communication efficiency. Compared with PS-based SGD algorithms, Rima also consumes less queue pairs (QPs) and can thus better leverage the power of RDMA/RoCE to accelerate the training speed. Our experiment shows that, compared with PS-based DSGD when solving 1M × 1M MF, Rima achieves comparable convergence performance after equal number of iterations, but reduces the training time by 68.7% and 85.4% via TCP and RDMA respectively. Jinkun Geng, Dan Li 0001, Shuai Wang 0028 |
ICDE | 2 |
| 2019 | DLBooster: Boosting End-to-End Deep Learning Workflows with Offloading Data Preprocessing PipelinesabstractIn recent years, deep learning (DL) has prospered again due to improvements in both computing and learning theory. Emerging studies mostly focus on the acceleration of refining DL models but ignore data preprocessing issues. However, data preprocessing can significantly affect the overall performance of end-to-end DL workflows. Our studies on several image DL workloads show that existing preprocessing backends are quite inefficient: they either perform poorly in throughput (30% degradation) or burn too many (>10) CPU cores. Based on these observations, we propose DLBooster, a high-performance data preprocessing pipeline that selectively offloads key workloads to FPGAs, to fit the stringent demands on data preprocessing for cutting-edge DL applications. Our testbed experiments show that, compared with the existing baselines, DLBooster can achieve 1.35×~2.4× image processing throughput in several DL workloads, but consumes only 1/10 CPU cores. Besides, it also reduces the latency by 1/3 in online image inference. Dan Li 0001, Binyao Jiang, Xi Fan, Jinkun Geng, Wei Bai 0001, Ran Shu 0001, Peng Cheng 0005, Yongqiang Xiong |
ICPP | 2 |
| 2019 | Impact of Network Topology on the Performance of DML: Theoretical Analysis and Practical FactorsabstractTo deal with the increasingly larger input data and model sizes, it has become necessary to scale the training of machine learning models to multiple nodes, even a server cluster, which we call distributed machine learning, or DML. However, DML utilizes more computation power at the cost of high communication overhead, which may limit the overall performance in turn. In this paper, we study the impact of network topology on the DML performance both in theory and in practice. We compare two representative network topologies, namely, Fat-Tree which is widely-used in modern data centers, and BCube, which is a low-cost and server-centric network topology, both running on top of RDMA. The results show that Fat-Tree not only has theoretically higher global synchronization time (GST) than BCube, but its practical GST (by NS-3 based simulation) is also considerably larger than the theoretical one. By analyzing the large-scale simulation traces, we find that the root cause for the gap in Fat-Tree comes from the load imbalance among the multiple parallel paths as well as the inevitable PFC frames, both of which do not appear in BCube. For a cluster of around 250 servers, BCube achieves 53%\sim 70% lower GST than Fat-Tree from the simulation. As a result, we suggest using server-centric network topology such as BCube, instead of the common Fat-Tree network, to build a special-purpose DML cluster, due to its parallel synchronization, RDMA friendliness, natural load balance, as well as low economical cost. Shuai Wang 0028, Dan Li 0001, Jinkun Geng |
INFOCOM | 2 |
| 2019 | Sphinx: A Transport Protocol for High-Speed and Lossy Mobile NetworksabstractModern mobile wireless networks have been demonstrated to be high-speed but lossy, while mobile applications have more strict requirements including reliability, goodput guarantee, bandwidth efficiency, and computation efficiency. Such a complicated combination of requirements and conditions in networks pushes the pressure to transport layer protocol design. We analyze and argue that few of existing network transport layer solutions are able to handle all these requirements. We design and implement Sphinx to satisfy the four requirements in high-speed and lossy networks. Sphinx has (1) a proactive coding-based method named semi-random LT codes for loss recovery, which estimates packet loss rate and adjusts the redundancy level accordingly, (2) a reactive retransmission method named Instantaneous Compensation Mechanism (ICM) for loss retransmission, which compensates the lost packets once actual loss exceeds the estimation, and (3) a parallel coding architecture, which leverages multi-core, shared memory and kernel-bypass DPDK. Prototype and evaluation show that Sphinx outperforms TCP and other coding solutions significantly in microbenchmarks across all four requirements, and improves the performance of applications such as video streaming and block data transfer. Dan Li 0001, Wenfei Wu, K. K. Ramakrishnan, Jinkun Geng, Fei Gui, Fanzhao Wang, Kai Zheng 0003 |
IPCCC | 2 |
| 2019 | Metro: An Efficient Traffic Fast Rerouting Scheme With Low OverheadabstractFailure is common instead of exception in large-scale networks. To provide high service quality to upper-layer applications, it is desired that a converged backup path can be rapidly launched when failure occurs. In this paper, we design an IP based Fast ReRouting (FRR) scheme called Metro, which can solve the traffic rerouting convergence problem after arbitrary single link/node failure with low stretch for the backup path. When failure occurs in the network, Metro first indicates all the network areas that would be affected by the failure, and then finds out a few bridge links to drain the traffic in the affected network area to the network area that is not affected by the failure. In this way, Metro does not configure tunnels, encapsulate or modify data packets, and hence it is easy to be deployed in current networks. Extensive simulations show that Metro can solve arbitrary single link/node failure with backup paths shorter than the state-of-the-art solutions, and about 98% of the backup path stretch in Metro are the same as the optimal tunnel scheme. Xuya Jia, Dan Li 0001, Jing Zhu 0007, Yong Jiang 0001 |
IEEE/ACM Trans. Netw. | 2 |
| 2018 | Dante: Enabling FOV-Aware Adaptive FEC Coding for 360-Degree Video StreamingabstractAs 360-degree videos grow dramatically in popularity, more applications demand the ability to stream 360-degree videos to wirelessly connected devices, such as smartphone headsets. However, the limited capacity and the unstable network conditions make wireless networks ill-suited to the requirements of 360-degree videos--high resolution and low delay. One common approach is to take advantage of the fact that the viewer only watches a small portion of the video around the field of view (FOV). This allows for better allocation of network bandwidth by prioritizing content the viewer actually watches. Previous efforts on 360-degree videos have largely focused on adapting the encoded bitrate to optimize video quality in the time-varying FOV. This paper follows the general FOV-aware approach but uses a different technique. Rather than adapting bitrate, we explore the opportunities of a custom underlying transport protocol for 360-degree videos. In particular, we make a case for using Forward Error Correction (FEC) coding over UDP to reduce video streaming delay (a key limitation of all TCP-based approaches). We present Dante, an FOV-aware UDP-based video streaming protocol that adapts to changing network conditions by dynamically choosing FEC redundancy levels based on how close the video content is to the FOV region. Experimental results show that Dante improves video quality (PSNR) by 20% to 30% over traditional UDP-based video streaming protocols and 40% over FOV-aware DASH. Zhetao Li, Fei Gui, Jinkun Geng, Dan Li 0001, Zhibo Wang 0001, Usama Zafar |
APNet | 4 |
| 2018 | BML: A High-performance, Low-cost Gradient Synchronization Algorithm for DML TrainingabstractIn distributed machine learning (DML), the network performance between machines significantly impacts the speed of iterative training. In this paper we propose BML, a new gradient synchronization algorithm with higher network performance and lower network cost than the current practice. BML runs on BCube network, instead of using the traditional Fat-Tree topology. BML algorithm is designed in such a way that, compared to the parameter server (PS) algorithm on a Fat-Tree network connecting the same number of server machines, BML achieves theoretically 1/k of the gradient synchronization time, with k/5 of switches (the typical number of k is 2∼4). Experiments of LeNet-5 and VGG-19 benchmarks on a testbed with 9 dual-GPU servers show that, BML reduces the job completion time of DML training by up to 56.4%. Dan Li 0001, Jinkun Geng, Yanshu Wang, Shuai Wang 0028, Shutao Xia |
NeurIPS | 2 |
| 2018 | Towards full virtualization of SDN infrastructure
Dan Li 0001, Yirong Yu, Jing Zhu 0007, Jinkun Geng |
Comput. Networks | 2 |
| 2018 | SVDC: A Highly Scalable Isolation Architecture for Virtualized Layer-2 Data Center NetworksabstractWhile large layer-2 networks are widely accepted as the network fabric for modern data centers and network virtualization is required to support multi-tenant cloud computing, existing network virtualization solutions are not specifically designed for layer-2 networks. In this paper, we designSVDC, a highly-scalable and low-overhead virtualization architecture for large layer-2 data center networks. By leveraging the emerging software defined networking (SDN) framework, SVDC decouples the global identifier of a virtual network from the identifier carried in the packet header. Hence, SVDC can scale to a great number of virtual networks with a very short tag in the packet header, which is never achieved by previous network virtualization solutions. SVDC enhances MAC-in-MAC encapsulation in a way that packets with overlapped MAC addresses are correctly forwarded even without in-packet global identifiers to differentiate the virtual networks they belong to. Besides, scalable and efficient layer-2 multicast and broadcast within virtual networks are also supported in SVDC. With extensive simulations and experiments, we show that SVDC is better than existing solutions in many aspects, particularly isolating virtual networks with high scalability and higher network goodput due to minimal packet header overhead. Congjie Chen, Dan Li 0001, Jun Li 0001, Konglin Zhu |
IEEE Trans. Cloud Comput. | 2 |
| 2018 | Dependency-Aware Data Locality for MapReduceabstractMapReduce effectively partitions and distributes computation workloads to a cluster of servers, facilitating today's big data processing. Given the massive data to be dispatched, and the intermediate results to be collected and aggregated, there have been a significant studies on data locality that seeks to co-locate computation with data, so as to reduce cross-server traffic in MapReduce. They generally assume that the input data have little dependency with each other, which however is not necessarily true for that of many real-world applications, and we show strong evidence that the finishing time of MapReduce tasks can be greatly prolonged with such data dependency. In this paper, we present Dependency-Aware Locality for MapReduce (DALM) for processing the real-world input data that can be highly skewed and dependent. DALM accommodates data-dependency in a data-locality framework, organically synthesizing the key components from data reorganization, replication, placement. Beside algorithmic design within the framework, we have also closely examined the deployment challenges, particularly in public virtualized cloud environments, and have implemented DALM on Hadoop 1.2.1 with Giraph 1.0.0. Its performance has been evaluated through both simulations and real-world experiments, and compared with that of state-of-the-art solutions. Xiaoqiang Ma, Xiaoyi Fan 0001, Jiangchuan Liu, Dan Li 0001 |
IEEE Trans. Cloud Comput. | 4 |
| 2017 | LOS: A High Performance and Compatible User-level Network Operating SystemabstractWith the ever growing speed of Ethernet NIC and more and more CPU cores on commodity X86 servers, the processing capability of the network stack in Linux kernel has become the bottleneck. Recently there is a trend on moving the network stack up to user level and bypassing the kernel. However, most of these stacks require changing the APIs or modifying the source code of applications, and hence are difficult to support legacy applications. In this work, we design and develop LOS, a user-level network operating system that not only gains high throughput and low latency by kernel-bypass technologies but also achieves compatibility with legacy applications. We successfully run Nginx and NetPIPE on top of LOS without touching the source code, and the experimental results show that LOS achieves significant throughput and latency gains compared with Linux kernel. Jinkun Geng, Du Lin, Ruilin Ling, Dan Li 0001 |
APNet | 7 |
| 2017 | Survivable and bandwidth-guaranteed embedding of virtual clusters in cloud data centersabstractCloud computing has emerged as a powerful and elastic platform for internet service hosting, yet it also draws concerns of the unpredictable performance of cloud-based services due to network congestion. To offer predictable performance, the virtual cluster abstraction of cloud services has been proposed, which enables allocation and performance isolation regarding both computing resources and network bandwidth in a simplified virtual network model. One issue arisen in virtual cluster allocation is the survivability of tenant services against physical failures. Existing works have studied virtual cluster backup provisioning with fixed primary embeddings, but have not considered the impact of primary embeddings on backup resource consumption. To address this issue, in this paper we study how to embed virtual clusters survivably in the cloud data center, by jointly optimizing primary and backup embeddings of the virtual clusters. We formally define the survivable virtual cluster embedding problem. We then propose a novel algorithm, which computes the most resource-efficient embedding given a tenant request. Since the optimal algorithm has high time complexity, we further propose a faster heuristic algorithm, which is several orders faster than the optimal solution, yet able to achieve similar performance. Besides theoretical analysis, we evaluate our algorithms via extensive simulations. Ruozhou Yu, Guoliang Xue, Xiang Zhang 0005, Dan Li 0001 |
INFOCOM | 4 |
| 2017 | Quick NAT: High performance NAT system on commodity platformsabstractNAT gateway is an important network system in today's IPv4 network when translating a private IPv4 address to a public address. However, traditional NAT system based on Linux Netfilter cannot achieve high network throughput to meet modern requirements such as data centers. To address this challenge, we improve the network performance of NAT system by three ways. First, we leverage DPDK to enable polling and zero-copy delivery, so as to reduce the cost of interrupt and packet copies. Second, we enable multiple CPU cores to process in parallel and use lock-free hash table to minimize the contention between CPU cores. Third, we use hash search instead of sequential search when looking up the NAT rule table. Evaluation shows that our Quick NAT system significantly improves the performance of NAT on commodity platforms. Dan Li 0001, Ruilin Ling |
LANMAN | 2 |
| 2017 | A survey of network update in SDN
Dan Li 0001, Konglin Zhu, Shutao Xia |
Frontiers Comput. Sci. | 1 |
| 2017 | Network Performance Aware Optimizations on IaaS CloudsabstractNetwork performance aware optimizations have long been a hot research topic to optimize distributed applications on traditional network environments. However, those optimization techniques rely on a few measurements on pair-wise network performance, and such direct use of network measurements is no longer valid on Infrastructure-as-a-service (IaaS) clouds. First, the direct calibration is ineffective. Network performance measurements may not represent the long-term performance (informally the stable component inside network performance) because of virtualization and network performance interference in the cloud. Second, the direct calibration is inefficient because the measurement overhead of all pair-wise link performance in a cluster becomes prohibitively high as the number of instances increases. To effectively and efficiently utilize existing network performance aware optimizations on IaaS clouds, we propose to reduce the measurement overhead and decouple the constant component from the dynamic network performance while minimizing the difference between the network performance and the constant component. For effectiveness, we use the constant component to guide the network performance aware optimizations. For efficiency, we exploit a non-negative matrix factorization (NMF) method to reduce the calibration overhead. Furthermore, we observe a tradeoff between effectiveness and efficiency, and develop an adaptive approach to capture this tradeoff. We demonstrate effectiveness and efficiency of our approach by adopting network performance aware optimizations on two kinds of basic applications, collective communications of MPI and generic topology mapping, and two real-world applications, namely N-body and conjugate gradient (CG). Our experiments on Amazon EC2 and simulations demonstrate significant calibration overhead reduction and performance improvement on guiding network performance aware optimizations, when comparing our approach to other state-of-the-art approaches. Yifan Gong 0003, Bingsheng He, Dan Li 0001 |
IEEE Trans. Computers | 3 |
| 2017 | 𝔽2 Tree: Rapid Failure Recovery for Routing in Production Data Center NetworksabstractFailures are not uncommon in production data center networks (DCNs) nowadays. It takes long time for the DCN routing to recover from a failure and find new forwarding paths, significantly impacting realtime and interactive applications at the upper layer. In this paper, we present a fault-tolerant DCN solution, called F2Tree, which is readily deployed in existing DNCs. F2Tree can significantly improve the failure recovery time only through a small amount of link rewiring and switch configuration changes. Through testbed and emulation experiments, we show that F2Tree can greatly reduce the routing recovery time after failure (by 78%) and improve the performance of upper layer applications when routing failure happens (96% less deadline-missing requests). Guo Chen 0001, Youjian Zhao, Hailiang Xu, Dan Pei, Dan Li 0001 |
IEEE/ACM Trans. Netw. | 5 |
| 2016 | DVMP: Incremental traffic-aware VM placement on heterogeneous servers in data centersabstractAs the tremendous momentum cloud computing has grown, the modern data center networks are facing challenge to handle the increasing traffic demand among virtual machines (VMs). Simply adding more switches and links may increase network capacity but at the same time increase the complexity and infrastructure cost. Thus, intelligent VM placement has been proposed to reduce the intra-DC traffic. Prior solutions model the traffic-aware VM placement problem as a Balanced Minimum K-cut Problem (BMKP). However, the assumptions of “once-for-all” VM placement on physical servers with equal VM slots are often not realistic in practical data centers, and thus the naive BMKP model may lead to suboptimal placement solutions. In this work, we revisit the problem by considering the server heterogeneity and propose an incremental traffic-aware VM placement algorithm. Given that the BMKP model cannot be directly applied, we make a number of transformations to re-establish the model. First, by introducing pseudo VM slots on physical servers with less VM slots, we allow the number of available VM slots of each server to be different. Second, pseudo edges with infinite costs are added between existing VMs, and thus previously deployed VMs on the same physical server will still be packed together. Third, a change on the number of pseudo VM slots is applied, so that existing VMs placed on different physical servers will still be separated. In this way, we reduce the problem to a new BMKP problem, which results in a much better solution. The evaluation results show that DVMP can reduce up to 28%, 39% and 55% traffic compared with naive BMKP model, greedy VM placement and random VM placement, respectively. Dan Li 0001, Syed Shah-e-Mardan Ali Rizvi, Fangxin Wang 0001, Wu He |
IWQoS | 1 |
| 2016 | PALS: Saving Network Power With Low Overhead to ISPs and ApplicationsabstractPower saving in the network infrastructure has received great attention in recent years. Power-aware traffic management is proposed in many works, in which a subset of routers/links are preferentially used to carry traffic while other links are activated only when traffic load is high. However, it remains challenging how to minimize the overhead to both ISPs and applications, which is important to the successful deployment of power-aware traffic management in a real network. This paper presents PALS, a new Power-Aware Link State routing based traffic management protocol. Compared with previous solutions, PALS remarkably reduces the overheads to ISPs and applications by the following innovations. First, PALS minimizes the forwarding table expansion due to dynamic power-aware routing, by using destination based routing instead of pairwise routing (e.g., MPLS). Second, PALS limits packet reordering for applications, by never splitting traffic between an IE (ingress-egress) router pair to multiple paths. Third, PALS significantly reduces the computation complexity of the power-aware routing algorithm, by running a simple path selection algorithm at each ingress router with the knowledge of local traffic information as well as global link utilization, which are much easier to obtain than global traffic matrix required by the state-of-the-art solutions (e.g., Zhang , IEEE ICNP 2010). Extensive simulations and testbed experiments show that, although bearing the simplicities to minimize the overhead, PALS saves satisfactory network power, with quick response to traffic variance and negligible impact on the packet delivery performance for applications. Dan Li 0001, Yirong Yu, Junxiao Shi, Beichuan Zhang 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2016 | CCDN: Content-Centric Data Center NetworksabstractData center networks continually seek higher network performance to meet the ever increasing application demand. Recently, researchers are exploring the method to enhance the data center network performance by intelligent caching and increasing the access points for hot data chunks. Motivated by this, we come up with a simple yet useful caching mechanism for generic data centers, i.e., a server caches a data chunk after an application on it reads the chunk from the file system, and then uses the cached chunk to serve subsequent chunk requests from nearby servers. To turn the basic idea above into a practical system and address the challenges behind it, we design content-centric data center networks (CCDNs), which exploits an innovative combination of content-based forwarding and location [Internet Protocol (IP)]-based forwarding in switches, to correctly locate the target server for a data chunk on a fully distributed basis. Furthermore, CCDN enhances traditional content-based forwarding to determine the nearest target server, and enhances traditional location (IP)-based forwarding to make high utilization of the precious memory space in switches. Extensive simulations based on real-world workloads and experiments on a test bed built with NetFPGA prototypes show that, even with a small portion of the server's storage as cache (e.g., 3%) and with a modest content forwarding information base size (e.g., 1000 entries) in switches, CCDN can improve the average throughput to get data chunks by 43% compared with a pure Hadoop File System (HDFS) system in a real data center. Dan Li 0001, Fangxin Wang 0001, Anke Li, K. K. Ramakrishnan, Ying Liu 0024, Xue (Steve) Liu |
IEEE/ACM Trans. Netw. | 2 |
| 2016 | DCloud: Deadline-Aware Resource Allocation for Cloud Computing JobsabstractWith the tremendous growth of cloud computing, it is increasingly critical to provide quantifiable performance to tenants and to improve resource utilization for the cloud provider. Though many recent proposals focus on guaranteeing job performance (with a particular note on network bandwidth) in the cloud, they usually lack efficient utilization of cloud resource, or vice versa. In this paper we present DCloud, which leverages the (soft) deadlines of cloud computing jobs to enable flexible and efficient resource utilization in data centers. With the deadline requirement of a job guaranteed, DCloud employs both time sliding (postponing the launching time of a job) and bandwidth scaling (adjusting the bandwidth associated with VMs) in resource allocation, so as to better match the resource allocated to the job with the cloud's residual resource. Extensive simulations and testbed experiments show that DCloud can accept much more jobs than existing solutions, and significantly increase the cloud provider's revenue with less cost for individual tenants. Dan Li 0001, Congjie Chen, Junjie Guan, Jing Zhu 0007, Ruozhou Yu |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | Rewiring 2 Links Is Enough: Accelerating Failure Recovery in Production Data Center NetworksabstractFailures are not uncommon in production data center networks (DCNs) nowadays, and it takes long time for the network to recover from a failure and find new forwarding paths, significantly impacting real time and interactive applications at the upper layer. The slow failure recovery is due to two primary reasons. First, there lacks immediate backup paths for downward links in DCN with multi-rooted tree topology. Second, distributed routing protocols in DCN take time to converge after failures. In this paper, we present a fault-tolerant DCN solution, called F2Tree, that can significantly improve the failure recovery time in current DCNs, only through a small amount of link rewiring and switch configuration changes. Because F2Tree does not change any existing software or hardware, it is readily deployed in production DCNs, where other existing proposals fail to achieve. Through testbed and emulation experiments, we show that F2Tree can greatly reduce the time of failure recovery by 78%. Our experimental results also show that, for partition-aggregate applications (popular in DCN) under various failure conditions, F2Tree reduces the ratio of deadline-missing requests by more than 96% compared to current DCNs. Guo Chen 0001, Youjian Zhao, Dan Pei, Dan Li 0001 |
ICDCS | 4 |
| 2015 | SVirt: A Substrate-agnostic SDN Virtualization Architecture for Multi-tenant CloudabstractData center operators are accepting software defined networking (SDN) to manage their networks, but it remains challenging how to provide desirable virtual SDN services to tenants in a public cloud. We design SVirt, which enables highly flexible virtual SDN in a multi-tenant cloud by a substrate-agnostic SDN virtualization architecture. By redesigning the physical switch's processing pipeline with a "late-binding key extractor", SVirt supports virtual SDN switches with different processing pipelines simultaneously on a physical switch. In the control plane, SVirt enables "many-to-one" and "one-to-many" mapping when allocating the physical resource for a virtual network, which embraces arbitrary topology and TCAM resource demanded by a virtual network. In the data plane, SVirt explicitly carries the forwarding context information in the packets, overcoming the "context-loss problem" in a virtual SDN network. We develop a NetFPGA prototype of SVirt switch. Evaluations based on event-driven simulations and prototype-based experiments demonstrate that, compared with traditional approaches, SVirt significantly enhances the cloud's capability to accept various virtual SDN requests and improves the network's throughput. Yirong Yu, Dan Li 0001 |
ICNP | 2 |
| 2015 | TAPS: Software Defined Task-Level Deadline-Aware Preemptive Flow Scheduling in Data CentersabstractMany data center applications have deadline requirements, which pose a requirement of deadline-awareness in network transport. Completing within deadlines is a necessary requirement for flows to be completed. Transport protocols in current data centers try to share the network resources fairly and are deadline-agnostic. Recently several works try to address the problem by making as many flows meet deadlines as possible. However, for many data center applications, a task cannot be completed until the last flow finishes, which indicates the bandwidths consumed by completed flows are wasted if some flows in the task cannot meet deadlines. In this paper we design a task-level deadline-aware preemptive flow scheduling(TAPS), which aims to make more tasks meet deadlines. We leverage software defined networking (SDN) technology and generalize SDN from flow-level awareness to task-level awareness. The scheduling algorithm runs on the SDN controller, which decides whether a flow should be accepted or discarded, pre-allocates the transmission time slices and computes the routing paths for accepted flows. Extensive flow-level simulations demonstrate TAPS outperforms Varys, Bara at, PDQ (Preemptive Distributed Quick flow scheduling), D3 (Deadline-Driven Delivery control protocol) and Fair Sharing transport protocols in deadline sensitive data center environment. A simple implementation on real systems also proves that TAPS makes high effective utilization of the network bandwidth in data centers. Dan Li 0001 |
ICPP | 2 |
| 2015 | MIFO: Multi-path Interdomain ForwardingabstractToday's interdomain routing is traffic agnostic when determining the single, best forwarding path. Naturally, as it does not adapt to congestion, the path chosen is not always optimal. In this paper, we focus on designing a multi-path interdomain forwarding (MIFO) mechanism, where AS border routers adaptively forward outbound traffic from a congested default path to an alternative path, without touching the interdomain routing protocols. Different from previous efforts which enable multi-path on control plane, MIFO achieves multi-path on data plane. The multiple alternative forwarding paths are obtained by exploring local BGP RIB. Multi-path forwarding on data plane can create a loop even within a stable network. MIFO solves this problem with a simple and practical approach. Several other challenges are also addressed including preventing cycling packet between iBGP peers and choosing the best alternative path from among multiple candidates. Our evaluations show that MIFO significantly improves the end-to-end throughput at the AS level, compared to traditional BGP and MIRO. For example, with only 50% of the ASes being MIFO capable, a significant percentage of the flows (about 40%) can use at least 50% of the inter-AS link capacity. In contrast, BGP and MIRO routing make less effective use of the inter-AS links, with only 7% and 17% of the flows can be so. Finally, we have developed a prototype implementation of MIFO on Linux with the forwarding engine in the kernel, with the routing daemon developed on XORP platform. The experiments on a test bed built with prototypes show that MIFO can improves the aggregate throughput by 81% compared with BGP routing. Dan Li 0001, Ying Liu 0024, Dan Pei, K. K. Ramakrishnan |
ICPP | 2 |
| 2015 | Rapier: Integrating routing and scheduling for coflow-aware data center networksabstractIn the data flow models of today's data center applications such as MapReduce, Spark and Dryad, multiple flows can comprise a coflow group semantically. Only completing all flows in a coflow is meaningful to an application. To optimize application performance, routing and scheduling must be jointly considered at the level of a coflow rather than individual flows. However, prior solutions have significant limitation: they only consider scheduling, which is insufficient. To this end, we present Rapier, a coflow-aware network optimization framework that seamlessly integrates routing and scheduling for better application performance. Using a small-scale testbed implementation and large-scale simulations, we demonstrate that Rapier significantly reduces the average coflow completion time (CCT) by up to 79.30% compared to the state-of-the-art scheduling-only solution, and it is readily implementable with existing commodity switches. Yangming Zhao, Kai Chen 0005, Wei Bai 0001, Minlan Yu, Chen Tian 0001, Yanhui Geng, Yiming Zhang 0003, Dan Li 0001, Sheng Wang 0006 |
INFOCOM | 8 |
| 2015 | Bandwidth guaranteed virtual network function placement and scaling in datacenter networksabstractEnterprises deploy their middlebox services in cloud seeking for easy management, flexible scalability and economic savings. However, existing elastic virtual network function(VNF) placement strategy often leads to an unpredictable placing location due to the ever-changing workload, which may waste much precious bandwidth resource and bring a lot of VM operation overhead(e.g. VM launch, termination and migration). A key problem for cloud providers is how to conduct an effective service placement and provide resource provision according to various workload, satisfying the bandwidth requirement of each service while saving as much cloud resource as possible. In this paper we solve both the virtual network function(VNF) placement and scaling problem based on preplanned allocation with bandwidth guarantee. We first propose a concept of VNF instance communication graph to describe the bandwidth demand of each VNF instance and explore the placement requirement for bandwidth savings. Then we design an on-line heuristic algorithm to achieve approximate optimal allocation. At last, we also provide an off-line optimal solution for comparison. Our simulation shows that our heuristic solution saves 20% more bandwidth resource and reduce more VM migration overhead than existing elastic placement solution. Its performance is also very close to the optimal solution. Fangxin Wang 0001, Ruilin Ling, Jing Zhu 0007, Dan Li 0001 |
IPCCC | 4 |
| 2015 | SIONA: A Service and Information Oriented Network Architecture
Mingwei Xu 0001, Zhongxing Ming, Chunmei Xia, Jia Ji, Dan Li 0001, Dan Wang 0002 |
J. Netw. Comput. Appl. | 5 |
| 2015 | On the Network Power Effectiveness of Data Center ArchitecturesabstractCloud computing not only requires high-capacity data center networks to accelerate bandwidth-hungry computations, but also causes considerable power expenses to cloud providers. In recent years many advanced data center network architectures have been proposed to increase the network throughput, such as Fat-Tree [1] and BCube [2], but little attention has been paid to the power efficiency of these network architectures. This paper makes the first comprehensive comparison study for typical data center networks with regard to their Network Power Effectiveness(NPE), which indicates the end-to-end bps per watt in data transmission and reflects the tradeoff between power consumption and network throughput. We take switches, server NICs and server CPU cores into account when evaluating the network power consumption. We measure NPE under both regular routing and power-aware routing, and investigate the impacts of topology size, traffic load, throughput threshold in power-aware routing, network power parameter as well as traffic pattern. The results show that in most cases Flattened Butterfly possesses the highest NPE among the architectures under study, and server-centric architectures usually have higher NPEs than Fat-Tree and VL2 architectures. In addition, the sleep-on-idle technique and power-aware routing can significantly improve the NPEs for all the data center architectures, especially when the traffic load is low. We believe that the results are useful for cloud providers, when they design/upgrade data center networks or employ network power management. Yunfei Shang, Dan Li 0001, Jing Zhu 0007, Mingwei Xu 0001 |
IEEE Trans. Computers | 2 |
| 2015 | Guaranteeing Heterogeneous Bandwidth Demand in Multitenant Data Center NetworksabstractThe ability to provide guaranteed network bandwidth for tenants is essential to the prosperity of cloud computing platforms, as it is a critical step for offering predictable performance to applications. Despite its importance, it is still an open problem for efficient network bandwidth sharing in a multitenant environment, especially when applications have diverse bandwidth requirements. More precisely, it is not only that different tenants have distinct demands, but also that one tenant may want to assign bandwidth differently across her virtual machines (VMs), i.e., the heterogeneous bandwidth requirements. In this paper, we tackle the problem of VM allocation with bandwidth guarantee in multitenant data center networks. We first propose an online VM allocation algorithm that improves on the accuracy of the existing work. Next, we develop a VM allocation algorithm under heterogeneous bandwidth demands. We conduct extensive simulations to demonstrate the efficiency of our method. Dan Li 0001, Jing Zhu 0007, Junjie Guan |
IEEE/ACM Trans. Netw. | 1 |
| 2015 | Willow: Saving Data Center Network Energy for Network-Limited FlowsabstractToday's giant data centers are power hungry. Data center energy saving not only helps control the operational cost, but also benefits the sustainable growth of cloud services. Due to the adoption of much more switches in modern data centers as well as the mature server-side power management techniques, energy saving for the data center network is becoming increasingly important. Most previous works on saving data center network energy focus on aggregating flows to as few switches as possible. However, in this paper we argue that this method may not work for network-limited flows, the throughputs of which are elastic based on the competing flows. To save the network energy consumed by this kind of elastic flows, we propose a flow scheduling approach called Willow, which takes both the number of switches involved and their active working durations into consideration. We formulate this problem by programming and design a greedy approximate algorithm to schedule flows in an online manner. Simulations based on MapReduce traces show that Willow can save up to 60 percent network energy compared with ECMP scheduling in typical settings, and outperforms other classical heuristic algorithms such as simulated annealing and particle swarm optimization. Testbed Experiments demonstrate that this kind of dynamic energy-efficient flow scheduling causes negligible impact on upper-layer applications. Dan Li 0001, Yirong Yu, Wu He, Kai Zheng 0003, Bingsheng He |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | Dependency-Aware Data Locality for MapReduceabstractRecent years have witnessed the prevalence of MapReduce-based systems, e.g., the Apache Hadoop, in large-scale distributed data processing. Fetching data from remote servers across multiple network switches is known to be costly. Hence, it is highly desirable to co-locate computation with data. State-of-the-art popularity-based replication achieves data locality through replicating popular files and spreading the replicas over multiple servers. While working well for independent files, they can store highly dependent files in different servers, resulting in excessive remote data accesses exchanges and consequently prolonging the job completion time. In this paper, we develop DALM (Dependency-Aware Locality for MapReduce), a novel replication strategy for general real-world input data that can be highly skewed and dependent. DALM accommodates data-dependency in a data-locality framework that comprehensively weights such key factors as popularity and storage budget. We extensively evaluate DALM through both simulations and real-world implementations, and have compared with state-of-the-art solutions, including the Hadoop system and the popularity-based Scarlett. The results show that DALM can significantly improve data locality for different inputs. For a popular iterative graph processing application on Hadoop, our prototype implementation of DALM reduces the remote data access and job completion time by 34.3% and 9.4%, respectively. Xiaoyi Fan 0001, Xiaoqiang Ma, Jiangchuan Liu, Dan Li 0001 |
IEEE CLOUD | 4 |
| 2014 | CDRDN: Content Driven Routing in Datacenter NetworkabstractA major challenge in data center networks is to provide enough network capacity to meet the ever increasing demand of large-scale distributed computing. While existing proposals focus on adding more switches and links, which cost extra hardware and energy, we explore another dimension in which spare disk space at servers is used for caching data, so as to increase network throughput without additional hardware. Leveraging the Named Data Networking (NDN) architecture and unique characteristics of data centers, we design a novel Content Driven Routing in Datacenter Network (CDRDN) that enables universal caching in data center networks in an efficient and scalable way. First, rather than using switches for “on-path” caching like in native NDN, CDRDN uses the large storage space at servers to do “off-path” caching. Second, by taking advantage of data center's regular and hierarchical network topology, CDRDN switches are able to direct requests to nearby server caches even under high dynamics of caches. Third, given the vast amount of data in data centers, the full name-based routing table would be difficult to fit in switch's limited fast memory. CDRDN adopts a compound content and location routing to ensure packet delivery while benefiting from name-based routing as much as the switches can afford. CDRDN extends NDN's adaptive forwarding mechanism to deal with cache misses, link failures, and congestion without running routing protocols or cache exchange protocols. Our packet-level simulations show that CDRDN can almost double the network throughput compared with shortest-path routing under the same setting, and CDRDN can effectively deal with link failures using adaptive forwarding. Dan Li 0001, Ying Liu 0024 |
ICCCN | 2 |
| 2014 | Freeway: Adaptively Isolating the Elephant and Mice Flows on Different Transmission PathsabstractThe network resource competition of today' data enters is extremely intense between long-lived elephant flows and latency-sensitive mice flows. Achieving both goals of high throughput and low latency respectively for the two types of flows requires compromise, which recent research has not successfully solved mainly due to the transfer of elephant and mice flows on shared links without any differentiation. However, current data enters usually adopt clos-based topology, e.g. Fat-tree/VL2, so there exist multiple shortest paths between any pair of source and destination. In this paper, we leverage on this observation to propose a flow scheduling scheme, Freeway, to adaptively partition the transmission paths into low latency paths and high throughput paths respectively for the two types of flows. An algorithm is proposed to dynamically adjust the number of the two types of paths according to the real-time traffic. And based on these separated transmission paths, we propose different flow type-specific scheduling and forwarding methods to make full utilization of the bandwidth. Our simulation results show that Freeway significantly reduces the delay of mice flow by 85.8% and achieves 9.2% higher throughput compared with Hedera. Wei Wang 0157, Yi Sun 0004, Kai Zheng 0003, Mohamed Ali Kâafar, Dan Li 0001, Zhongcheng Li |
ICNP | 5 |
| 2014 | TED: Inter-domain traffic engineering via deflectionabstractAs inter-domain routing on today's Internet does not and basically cannot consider traffic load when determining best traffic forwarding paths, it is not always optimal for a router to forward packets along its default path, especially when the router's default output port incurs a long queuing delay. In this paper, we design a new approach called TED in which border routers of autonomous systems (AS) adaptively deflect outbound traffic from a congested default path to an alternative path to significantly improve inter-domain traffic engineering (TE) and end-to-end throughput. With TED, every router only needs to examine the queue length of its own outgoing ports to orchestrate its deflection operation and ensure traffic forwarding is at line speed. It does not need to communicate or coordinate with other TED-capable routers or modify packet content, making TED incrementally deployable. Our evaluation shows that TED significantly increases the average throughput of traffic flows, and the improvement is comparable to directly upgrading router hardware and capacity. Finally, a prototype of TED on NetFPGA is also implemented. Jun Li 0001, Ying Liu 0024, Dan Li 0001 |
IWQoS | 4 |
| 2014 | Finding Constant from Change: Revisiting Network Performance Aware Optimizations on IaaS CloudsabstractNetwork performance aware optimizations have long been an effective approach to optimizing distributed applications on traditional network environments. However, the assumptions of network topology or direct use of several measurements of pair-wise network performance for optimizations are no longer valid on IaaS clouds. Virtualization hides network topology from users, and direct use of network performance measurements may not represent long-term performance. To enable existing network performance aware optimizations on IaaS clouds, we propose to decouple constant component from dynamic network performance while minimizing the difference by a mathematical method called RPCA (Robust Principal Component Analysis). We use the constant component to guide network performance aware optimizations and demonstrate the efficiency of our approach by adopting network aware optimizations for collective communications of MPI and generic topology mapping as well as two real-world applications, N-body and conjugate gradient (CG). Our experiments on Amazon EC2 and simulations demonstrate significant performance improvement on guiding the optimizations. Yifan Gong 0003, Bingsheng He, Dan Li 0001 |
SC | 3 |
| 2014 | LTTP: An LT-Code Based Transport Protocol for Many-to-One Communication in Data CentersabstractTCP has been widely adopted in current data centers to ensure reliable data delivery. However, recently TCP Incast was found to occur in many-to-one communications with barrier-synchronized requirement, where the TCP goodput drops dramatically. Previous solutions to TCP Incast either require updating the OS/hardware to support fine-grained timers, or smartly control utilization of the switch buffer to reduce the probability of buffer overflow and packet loss. In this paper we explore a different approach to support many-to-one communication in data center networks, which we call LTTP (LT-code based Transport Protocol). LTTP improves LT (Luby Transform) code to achieve reliable UDP-based transmission by exploiting data redundancy, and employs TFRC (TCP Friendly Rate Control) to adjust the traffic sending rates at servers. NS-2 based simulation shows that the goodput of LTTP never degrades with the increase of the number of servers in many-to-one communications, and LTTP significantly outperforms DCTCP when the number of servers is large. Simulation results also demonstrate that LTTP flows can fairly share bandwidth with TCP flows. Changlin Jiang, Dan Li 0001, Mingwei Xu 0001 |
IEEE J. Sel. Areas Commun. | 2 |
| 2014 | GreenDCN: A General Framework for Achieving Energy Efficiency in Data Center NetworksabstractThe popularization of cloud computing has raised concerns over the energy consumption that takes place in data centers. In addition to the energy consumed by servers, the energy consumed by large numbers of network devices emerges as a significant problem. Existing work on energy-efficient data center networking primarily focuses on traffic engineering, which is usually adapted from traditional networks. We propose a new framework to embrace the new opportunities brought by combining some special features of data centers with traffic engineering. Based on this framework, we characterize the problem of achieving energy efficiency with a time-aware model, and we prove its NP-hardness with a solution that has two steps. First, we solve the problem of assigning virtual machines (VM) to servers to reduce the amount of traffic and to generate favorable conditions for traffic engineering. The solution reached for this problem is based on three essential principles that we propose. Second, we reduce the number of active switches and balance traffic flows, depending on the relation between power consumption and routing, to achieve energy conservation. Experimental results confirm that, by using this framework, we can achieve up to 50 percent energy savings. We also provide a comprehensive discussion on the scalability and practicability of the framework. Lin Wang 0015, Fa Zhang 0001, Jordi Arjona Aroca, Athanasios V. Vasilakos, Kai Zheng 0003, Chenying Hou, Dan Li 0001, Zhiyong Liu 0002 |
IEEE J. Sel. Areas Commun. | 7 |
| 2014 | Reliable Multicast in Data Center NetworksabstractMulticast benefits data center group communication in both saving network traffic and improving application throughput. Reliable packet delivery is required in data center multicast for data-intensive computations. However, existing reliable multicast solutions for the Internet are not suitable for the data center environment, especially with regard to keeping multicast throughput from degrading upon packet loss, which is norm instead of exception in data centers. We present RDCM, a novel reliable multicast protocol for data center network. The key idea of RDCM is to minimize the impact of packet loss on the multicast throughput, by leveraging the rich link resource in data centers. A multicast-tree-aware backup overlay is explicitly built on group members for peer-to-peer packet repair. The backup overlay is organized in such a way that it causes little individual repair burden, control overhead, as well as overall repair traffic. RDCM also realizes a window-based congestion control to adapt its sending rate to the traffic status in the network. Simulation results in typical data center networks show that RDCM can achieve higher application throughput and less traffic footprint than other representative reliable multicast protocols. We have implemented RDCM as a user-level library on Windows platform. The experiments on our test bed show that RDCM handles packet loss without obvious throughput degradation during high-speed data transmission, gracefully respond to link failure and receiver failure, and causes less than 10% CPU overhead to data center servers. Dan Li 0001, Mingwei Xu 0001, Ying Liu 0024, Yong Cui 0001, Guihai Chen |
IEEE Trans. Computers | 1 |
| 2014 | Revisiting the Design of Mega Data Centers: Considering Heterogeneity Among ContainersabstractIn this paper, we revisit the design of mega data centers, which are usually built by a number of modularized containers. Due to technical innovation and vendor diversity, heterogeneity widely exists among data-center containers in practice. To embrace this issue, we propose uFix, which is a scalable, flexible, and modularized network architecture to interconnect heterogeneous data-center containers. The intercontainer connection rule in uFix is designed in such a way that it can flexibly scale to a huge number of servers with stable server/switch hardware settings. uFix allows modularized and fault-tolerant routing by completely decoupling intercontainer routing from intracontainer routing. We implement a software-based uFix prototype on a Linux platform. Both simulation and prototype-based experiment show that uFix enjoys high network capacity, gracefully handles server/switch failures, and causes lightweight CPU overhead onto data-center servers. Dan Li 0001, Mingwei Xu 0001, Xiaoming Fu 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2013 | PACE: Policy-Aware Application Cloud EmbeddingabstractThe emergence of new capabilities such as virtualization and elastic (private or public) cloud computing infrastructures has made it possible to deploy multiple applications, on demand, on the same cloud infrastructure. A major challenge to achieve this possibility, however, is that modern applications are typically distributed, structured systems that include not only computational and storage entities, but also policy entities (e.g., load balancers, firewalls, intrusion prevention boxes). Deploying applications on a cloud infrastructure without the policy entities may introduce substantial policy violations and/or security holes. In this paper, we present PACE: the first systematic framework for Policy-Aware Application Cloud Embedding. We precisely define the policy-aware, cloud application embedding problem, study its complexity and introduce simple, efficient, online primal-dual algorithms to embed applications in cloud data centers. We conduct evaluations using data from a real, large campus network and a realistic data center topology to evaluate the feasibility and performance of PACE. We show that deployment in a cloud without considering in-network policies may lead to a large number of policy violations (e.g., using tree routing as a way to enforce in-network policies may observe up to 91% policy violations). We also show that our embedding algorithms are very efficient by comparing with a good online fractional embedding algorithm. Li Erran Li, Vahid Liaghat, Mohammad Hajiaghayi, Dan Li 0001, Gordon T. Wilfong, Yang Richard Yang, Chuanxiong Guo |
INFOCOM | 5 |
| 2013 | Resource management in radio access and IP-based core networks for IMT Advanced and Beyond
Gang Su, Markus Hidell, Henrik Abrahamsson, Bengt Ahlgren, Dan Li 0001, Peter Sjödin, Voravit Tanyingyong, Ke Xu 0002 |
Sci. China Inf. Sci. | 5 |
| 2013 | Expandable and Cost-Effective Network Structures for Data Centers Using Dual-Port ServersabstractA fundamental goal of data center networking is to efficiently interconnect a large number of servers with the low equipment cost. Several server-centric network structures for data centers have been proposed. They, however, are not truly expandable and suffer a low degree of regularity and symmetry. Inspired by the commodity servers in today's data centers that come with dual port, we consider how to build expandable and cost-effective structures without expensive high-end switches and additional hardware on servers except the two NIC ports. In this paper, two such network structures, called HCN and BCN, are designed, both of which are of server degree 2. We also develop the low overhead and robust routing mechanisms for HCN and BCN. Although the server degree is only 2, HCN can be expanded very easily to encompass hundreds of thousands servers with the low diameter and high bisection width. Additionally, HCN offers a high degree of regularity, scalability, and symmetry, which conform to the modular designs of data centers. BCN is the largest known network structure for data centers with the server degree 2 and network diameter 7. Furthermore, BCN has many attractive features, including the low diameter, high bisection width, large number of node-disjoint paths for the one-to-one traffic, and good fault-tolerant ability. Mathematical analysis and comprehensive simulations show that HCN and BCN possess excellent topological properties and are viable network structures for data centers. Deke Guo, Tao Chen 0013, Dan Li 0001, Mo Li 0001, Yunhao Liu 0001, Guihai Chen |
IEEE Trans. Computers | 3 |
| 2013 | Dynamic Scheduling for Wireless Data Center NetworksabstractUnbalanced traffic demands of different data center applications are an important issue in designing data center networks (DCN). In this paper, we present our exploratory investigation on a hybrid DCN solution of utilizing wireless transmissions in DCNs. Our work aims to solve the congestion problem caused by a few hot nodes to improve the global performance. We model the wireless transmissions in DCN by considering both the wireless interference and the adaptive transmission rate. Besides, both throughput and job completion time are considered to measure the impact of wireless transmissions on the global performance. Based on the model, we formulate the problem of channel allocation as an optimization problem. We also design an approximation algorithm with an approximation bound of 1/2 and a genetic algorithm to address the scheduling problem. A series of simulations are performed to evaluate and demonstrate the effectiveness of our wireless DCN scheme. Yong Cui 0001, Hongyi Wang 0004, Xiuzhen Cheng, Dan Li 0001, Antti Ylä-Jääski |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2013 | IP-Geolocation Mapping for Moderately Connected Internet RegionsabstractMost IP-geolocation mapping schemes [14], [16], [17], [18] take delay-measurement approach, based on the assumption of a strong correlation between networking delay and geographical distance between the targeted client and the landmarks. In this paper, however, we investigate a large region of moderately connected Internet and find the delay-distance correlation is weak. But we discover a more probable rule - with high probability the shortest delay comes from the closest distance. Based on this closest-shortest rule, we develop a simple and novel IP-geolocation mapping scheme for moderately connected Internet regions, called GeoGet. In GeoGet, we take a large number of webservers as passive landmarks and map a targeted client to the geolocation of the landmark that has the shortest delay. We further use JavaScript at targeted clients to generate HTTP/Get probing for delay measurement. To control the measurement cost, we adopt a multistep probing method to refine the geolocation of a targeted client, finally to city level. The evaluation results show that when probing about 100 landmarks, GeoGet correctly maps 35.4 percent clients to city level, which outperforms current schemes such as GeoLim [16] and GeoPing [14] by 270 and 239 percent, respectively, and the median error distance in GeoGet is around 120 km, outperforming GeoLim and GeoPing by 37 and 70 percent, respectively. Dan Li 0001, Chuanxiong Guo, Yunxin Liu 0001, Zhi-Li Zhang, Yongguang Zhang |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2012 | SIONA: A service and information oriented network architectureabstractThe Internet is a great hit in human history. However, it has evolved greatly from its original incarnation. Content distribution is playing a central role in today's Internet, which makes it difficult for the conventional host-to-host communication to meet the ever-increasing demands. In this paper, we present a novel “service and information oriented network architecture” (SIONA). The key aspect of SIONA is the name-based two-dimensional routing paradigm that provides scalable routing, caching and content delivery. We argue that SIONA solves the problems of mobile Internet by naturally supporting mobility, and provides network layer P2P for massive data distribution. Evaluation is conducted to investigate its caching and mobility performance. Zhongxing Ming, Mingwei Xu 0001, Chunmei Xia, Dan Li 0001, Dan Wang 0002 |
ICC | 4 |
| 2012 | Towards bandwidth guarantee in multi-tenancy cloud computing networksabstractTo efficiently utilize their infrastructure and thus increase their revenue, cloud providers need mechanisms to provide resource allocation and performance isolation for different tenants in the shared platform. In particular, network bandwidth sharing is a critical yet still an open problem to most cloud providers. In this paper, we study the problem of virtual machine (VM) allocation under the consideration of providing bandwidth guarantees. We first propose an online allocation algorithm for tenants with homogeneous bandwidth demand, which improves on the accuracy of existing algorithms. Subsequently, we extend it to handle heterogeneous bandwidth demand. Extensive simulations show that our algorithm makes much more efficient utilization of the network resource than existing algorithms, and performs close to the optimal offline allocation. Jing Zhu 0007, Dan Li 0001, Hongnan Liu |
ICNP | 2 |
| 2012 | ESM: Efficient and Scalable Data Center Multicast RoutingabstractMulticast benefits group communications in saving network traffic and improving application throughput, both of which are important for data center applications. However, the technical trend of data center design poses new challenges for efficient and scalable multicast routing. First, the densely connected networks make traditional receiver-driven multicast routing protocols inefficient in multicast tree formation. Second, it is quite difficult for the low-end switches widely used in data centers to hold the routing entries of massive multicast groups. In this paper, we propose ESM, an efficient and scalable multicast routing scheme for data center networks. ESM addresses the challenges above by exploiting the feature of modern data center networks. Based on the regular topology of data centers, ESM uses a source-to-receiver expansion approach to build efficient multicast trees, excluding many unnecessary intermediate switches used in receiver-driven multicast routing. For scalable multicast routing, ESM combines both in-packet Bloom Filters and in-switch entries to make the tradeoff between the number of multicast groups supported and the additional bandwidth overhead. Simulations show that ESM saves 40%$\sim$50% network traffic and doubles the application throughputs compared to receiver-driven multicast routing, and the combination routing scheme significantly reduces the number of in-switch entries required. We implement ESM on a Linux platform. The experimental results further demonstrate that ESM can well support online tree building for large-scale groups with churns, and the overhead of the combination forwarding engine is light-weighted. Dan Li 0001, Yuanjie Li, Sen Su, Jiangwei Yu |
IEEE/ACM Trans. Netw. | 1 |
| 2011 | Building mega data center from heterogeneous containersabstractData center containers are regarded as the basic units to build mega data centers. In practice, heterogeneity exists among data center containers, because of technical innovation and vendor diversity. In this paper, we propose uFix, a scalable, flexible and modularized network architecture to interconnect heterogeneous data center containers. The inter-container connection rule in uFix is designed in such a way that it can flexibly scale to a huge number of servers with stable server/switch hardware settings. uFix allows modularized and fault-tolerant routing by completely decoupling inter-container routing from intra-container routing. We implement a software-based uFix stack on the Linux platform. Simulation and experiment results show that uFix enjoys high network capacity, gracefully handles server/switch failures, and brings light-weight CPU overhead onto data center servers. Dan Li 0001, Mingwei Xu 0001, Xiaoming Fu 0001 |
ICNP | 1 |
| 2011 | BCN: Expansible network structures for data centers using hierarchical compound graphsabstractA fundamental challenge in data centers is how to design networking structures for efficiently interconnecting a large number of servers. Several server-centric structures have been proposed, but are not truly expansible and suffer low degree of regularity and symmetry. To address this issue, we propose two novel structures called HCN and BCN, which utilize hierarchical compound graphs to interconnect large population of servers each with two ports only. They own two topological advantages, i.e., the expansibility and equal degree. In addition, HCN offers high degree of regularity, scalability and symmetry, which well conform to the modular design of data centers. Moreover, a BCN of level one in each dimension involves more servers than FiConn with server degree 2 and diameter 7, and is large enough for a single data center. Mathematical analysis and comprehensive simulations show that BCN possesses excellent topology properties and is a viable network structure for data centers. Deke Guo, Tao Chen 0013, Dan Li 0001, Yunhao Liu 0001, Xue (Steve) Liu, Guihai Chen |
INFOCOM | 3 |
| 2011 | RDCM: Reliable data center multicastabstractMulticast benefits data center group communication in both saving network traffic and improving application throughput. The SLA (Service Level Agreement) of cloud service requires the computation correctness of distributed applications, translating to the requirement of reliable Multicast delivery. In this paper we present RDCM, a novel reliable Multicast approach for data center network. The key idea of RDCM is to minimize the impact of packet loss on the Multicast performance, by leveraging the rich link resource in data centers. A Multicast-tree-aware backup overlay is purposely built on group members for peer-to-peer packet repair. Riding on Unicast, packet repair not only achieves complete repair isolation, but also has high probability to bypass the pathological links in the Multicast tree where packet loss occurs. The backup overlay is organized in such a way that it causes little individual repair burden, control overhead, as well as overall repair traffic. We have implemented RDCM as a user-level library on Windows platform. The experiments on our test bed show that RDCM handles packet loss without obvious throughput degradation during high-speed data transmission. Dan Li 0001, Mingwei Xu 0001, Ming-Chen Zhao, Chuanxiong Guo, Yongguang Zhang, Min-You Wu |
INFOCOM | 1 |
| 2011 | Exploring efficient and scalable multicast routing in future data center networksabstractMulticast benefits group communications in saving network traffic and improving application throughput, both of which are important for data center applications. However, the technical trend of future data center design poses new challenges for efficient and scalable Multicast routing. First, the densely connected networks make traditional receiver-driven Multicast routing protocols inefficient in Multicast tree formation. Second, it is quite difficult for the low-end switches largely used in data centers to hold the routing entries of massive Multicast groups. Dan Li 0001, Jiangwei Yu, Junbiao Yu |
INFOCOM | 1 |
| 2011 | Impact of user selfishness in construction action on the streaming quality of overlay multicast
Dan Li 0001, Yong Cui 0001, Jiangchuan Liu, Ke Xu 0002 |
Comput. Networks | 1 |
| 2011 | Scalable and cost-effective interconnection of data-center servers using dual server portsabstractThe goal of data-center networking is to interconnect a large number of server machines with low equipment cost while providing high network capacity and high bisection width. It is well understood that the current practice where servers are connected by a tree hierarchy of network switches cannot meet these requirements. In this paper, we explore a new server-interconnection structure. We observe that the commodity server machines used in today's data centers usually come with two built-in Ethernet ports, one for network connection and the other left for backup purposes. We believe that if both ports are actively used in network connections, we can build a scalable, cost-effective interconnection structure without either the expensive higher-level large switches or any additional hardware on servers. We design such a networking structure called FiConn. Although the server node degree is only 2 in this structure, we have proven that FiConn is highly scalable to encompass hundreds of thousands of servers with low diameter and high bisection width. We have developed a low-overhead traffic-aware routing mechanism to improve effective link utilization based on dynamic traffic state. We have also proposed how to incrementally deploy FiConn. Dan Li 0001, Chuanxiong Guo, Kun Tan 0001, Yongguang Zhang, Songwu Lu |
IEEE/ACM Trans. Netw. | 1 |
| 2011 | Defending Against Distance Cheating in Link-Weighted Application-Layer MulticastabstractApplication-layer multicast (ALM) has recently emerged as a promising solution for diverse group-oriented applications. Unlike dedicated routers in IP multicast, the autonomous end-hosts are generally unreliable and even selfish. A strategic host might cheat about its private information to affect protocol execution and, in turn, to improve its individual benefit. Specifically, in a link-weighted ALM protocol where the hosts measure the distances from their neighbors and accordingly construct the ALM topology, a selfish end-host can easily intercept the measurement message and exaggerate the distances to other nodes, so as to reduce the probability of being a relay. Such distance cheating, rarely happening in IP multicast, can significantly impact the efficiency and stability of the ALM topology. To defend against this kind of cheating, we present a Vickrey–Clarke–Groves (VCG)-based cheat-proof mechanism in this paper. We demonstrate a practical mapping from the utility, payment, and welfare of a VCG mechanism to the link-weighted ALM context. Based on this, we further discuss practical issues for implementing the cheat-proof mechanism—specifically, a trustworthy distributed algorithm for payment computation. Performance analyses show that the overheads of the computation, storage, and communication of our implementation are controlled at low levels, and extensive simulations further testify the implementation's effectiveness. Although there are other similar studies in this area, the contribution of our cheat-proof mechanism and its implementation primarily lies in two aspects. On one hand, we first explicitly solve the distance cheating problem in link-weighted ALM since its proposal by mapping the VCG mechanism to link-weighted ALM context. On the other hand, our distributed implementation can not only effectively defend against distance cheating, but can also avoid the potential cheating behaviors when selfish ALM nodes fulfill the cheat-proof mechanism itself. Dan Li 0001, Jiangchuan Liu, Yong Cui 0001, Ke Xu 0002 |
IEEE/ACM Trans. Netw. | 1 |
| 2010 | WIND: A scalable and lightweight network topology service for peer-to-peer applicationsabstractWe present an Internet-scale network topology information (NTI) service named WIND for localizing P2P traffic. Central to WIND are the two simple ideas: 1) obtaining NTI directly from routing infrastructures, and 2) leveraging existing, widely deployed DNS caches for NTI delivery. WIND fulfills the fidelity, flexibility and scalability requirement of an effective NTI service. WIND is deployed in CERNET. We conduct extensive trace-driven emulations on PlanetLab. Experimental results confirm the effectiveness of the WIND service. Hongqiang Liu, Yongqiang Xiong, CongXiao Bao, Xing Li 0001, Guobin Shen, Dan Li 0001 |
NOMS | 6 |
| 2009 | MDCube: a high performance network structure for modular data center interconnectionabstractShipping-container-based data centers have been introduced as building blocks for constructing mega-data centers. However, it is a challenge on how to interconnect those containers together with reasonable cost and cabling complexity, due to the fact that a mega-data center can have hundreds or even thousands of containers and the aggregate bandwidth among containers can easily reach tera-bit per second. As a new inner-container server-centric network architecture, BCube [9] interconnects thousands of servers inside a container and provides high bandwidth support for typical traffic patterns. It naturally serves as a building block for mega-data center. Guohan Lu, Dan Li 0001, Chuanxiong Guo, Yongguang Zhang |
CoNEXT | 3 |
| 2009 | FiConn: Using Backup Port for Server Interconnection in Data CentersabstractThe goal of data center networking is to interconnect a large number of server machines with low equipment cost, high and balanced network capacity, and robustness to link/server faults. It is well understood that, the current practice where servers are connected by a tree hierarchy of network switches cannot meet these requirements (Fares et al., 2008 and Guo et al., 2008). In this paper, we explore a new server-interconnection structure. We observe that the commodity server machines used in today's data centers usually come with two built-in Ethernet ports, one for network connection and the other left for backup purpose. We believe that, if both ports are actively used in network connections, we can build a low-cost interconnection structure without the expensive higher-level large switches. Our new network design, called FiConn, utilizes both ports and only the low-end commodity switches to form a scalable and highly effective structure. Although the server node degree is only two in this structure, we have proven that FiConn is highly scalable to encompass hundreds of thousands of servers with low diameter and high bisection width. The routing mechanism in FiConn balances different levels of links. We have further developed a low-overhead traffic-aware routing mechanism to improve effective link utilization based on dynamic traffic state. Simulation results have demonstrated that the routing mechanisms indeed achieve high networking throughput. Dan Li 0001, Chuanxiong Guo, Kun Tan 0001, Songwu Lu |
INFOCOM | 1 |
| 2009 | BCube: a high performance, server-centric network architecture for modular data centersabstractThis paper presents BCube, a new network architecture specifically designed for shipping-container based, modular data centers. At the core of the BCube architecture is its server-centric network structure, where servers with multiple network ports connect to multiple layers of COTS (commodity off-the-shelf) mini-switches. Servers act as not only end hosts, but also relay nodes for each other. BCube supports various bandwidth-intensive applications by speeding-up one-to-one, one-to-several, and one-to-all traffic patterns, and by providing high network capacity for all-to-all traffic. Chuanxiong Guo, Guohan Lu, Dan Li 0001, Yunfeng Shi, Chen Tian 0001, Yongguang Zhang, Songwu Lu |
SIGCOMM | 3 |
| 2009 | Defending Against Buffer Map Cheating in DONet-Like P2P StreamingabstractData-driven overlay network (DONet)-like P2P system is especially suitable to support live stream applications, since its data structure can tolerate node dynamics quite well. However, optimal streaming demands the cooperation of individual nodes. If selfish nodes cheat about their buffer maps to reduce the forwarding burden, the overall streaming quality would be negatively affected. To defend against this kind of cheating, we design a trustworthy service-differentiation based incentive mechanism with low complexity in this paper. The mechanism is composed of the service-differentiation algorithm and the contribution-evaluation algorithm. Compared with other studies in this area, the primary characteristic of our mechanism lies in two aspects. Firstly, the contribution of each node is evaluated considering the features of live streaming, not just by the transferring bytes. Secondly, the potential cheating behavior of overlay nodes during the fulfillment of incentive algorithms can be avoided, which is usually not considered by other similar studies. Extensive simulations suggest that the algorithms are indeed effective for defending against buffer map cheating in DONet-like P2P streaming. Dan Li 0001, Yong Cui 0001 |
IEEE Trans. Multim. | 1 |
| 2007 | Truthful Streaming in Selfish DONetabstractData-driven overlay network (DONet) is especially suitable for live stream because it can tolerant node dynamics well. However, optimal streaming demands the cooperation of individual nodes. If selfish nodes in DONet cheat about their buffer maps to reduce the forwarding burden, the overall streaming quality might be negatively affected. To defend this kind of cheating behavior, we design a trustworthy service- differentiation based incentive mechanism with low complexity in this paper. The mechanism is composed of the service- differentiation algorithm and the contribution-evaluation algorithm. Compared with other studies in this area, the primary characteristic of our mechanism lies in two aspects. Firstly, the contribution of each node is evaluated considering the characteristic of live stream, not just by the transferring bytes. Secondly, the potential cheating behavior of overlay nodes during the fulfillment of incentive algorithms can be defended, which is usually not considered by other studies. Dan Li 0001, Yong Cui 0001 |
ICC | 1 |
| 2007 | QoS-Aware Streaming in Overlay Multicast Considering the Selfishness in Construction ActionabstractMost existing overlay multicast proposals have assumed that the nodes are cooperative and thus focus on the global topology optimization. However, a unique and important characteristic of overlay nodes is that, as application-layer agents, they can be selfish with their own interests. To achieve better quality-of-service (QoS) or to minimize forwarding overhead, an overlay node can behave selfishly in the information collection or in the overlay construction. While the former has recently been investigated, the impact of selfishness in the construction action remains unclear. In this paper, we present the first systematic study on the impact of selfishness in both tree and mesh overlay construction. Our investigation considers multiple QoS measures for streaming applications, including stream latency, resolution, and continuity. Our contribution is twofold: first, we analyze how for selfish overlay nodes to choose a construction-action policy to optimize their individual multi-metric QoS. Second, we demonstrate that the selfishness-aware policy for the construction action is consistent with the QoS optimization for the global multicast session, but not vice versa. The implication is significant: A globally optimal overlay construction itself can be vulnerable to individual selfishness; but, following our directions, we can design an overlay that is both globally optimal and selfish-resistant. Dan Li 0001, Yong Cui 0001, Jiangchuan Liu |
INFOCOM | 1 |
| 2006 | Segment-sending Schedule in Data-driven Overlay NetworkabstractData-driven overlay network is suitable for live-event streaming, because it can provide relatively-continuous streaming even in dynamic environment. In terms of improving streaming quality, prior work covered membership management, buffer map exchange, segment requesting schedule, etc. In this paper, we address the problem of segment-sending schedule on the segment-providing node, which may also affect the streaming quality. The schedule methods we discuss include FIFO schedule, lower-sequence favored schedule, and higher-sequence favored schedule. Simulation results show that if users care playing continuity much more than playing delay, the higher-sequence favored schedule brings the best streaming quality; however, if users care playing delay much more than playing continuity, lower-sequence favored schedule is the preferred choice. Through this work, we find another way to improve streaming quality in data-driven overlay network. Dan Li 0001, Yong Cui 0001, Ke Xu 0002 |
ICC | 1 |
| 2006 | Stability of ALM Tree with selfish receivers: A simulation study
Dan Li 0001, Yong Cui 0001 |
Comput. Commun. | 2 |
| 2005 | Impact of receiver cheating on the stability of ALM treeabstractApplication layer multicast (ALM) is an effective supplement to IP multicast, but it has the potential trouble of trust on end systems. For instance, multicast receivers may cheat in order to obtain a better position in the multicast tree. Receiver cheating may transform the multicast tree, and lead to its instability. We establish the cheating model of ALM receivers and analyze the stability of ALM tree when receiver cheating occurs. Simulation results show that receiver cheating has considerably negative effects on the stability of ALM tree. This discovery brings forward an issue in ALM study, that is, we should take receiver cheating into consideration to maintain a stable ALM tree when designing ALM protocols. Dan Li 0001, Yong Cui 0001, Ke Xu 0002 |
GLOBECOM | 1 |