EDBT 2026 Demo / reviewers in the wild / expert
Danfeng Shan
dblp:137/3349
· DBLP profile ↗
37ranked-venue papers
13as first author
24since 2021 · last 2026
0000-0003-0852-5955ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 31 · 10 first-author · 19 since 2021Systems, architecture and hardware · 4 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Quanta: Scaling Packet-Level Network Simulation by Exploiting Execution RedundancyabstractPacket-level network simulation provides high-fidelity modeling but suffers from severe scalability bottlenecks. Existing scaling approaches remain inefficient for modern data-center and AI-training networks. Spatial parallelism requires substantial hardware resources, while temporal-skipping approaches become less effective under bursty traffic. We observe that homogeneous data-center deployments introduce substantial execution redundancy during simulation. Jiajun Luan, Hao Li 0011, Yihan Dang, Ze Xia, Danfeng Shan, Peng Zhang 0011 |
APNet | 5 |
| 2026 | CCC: Re-architecting Delay-based Congestion Control in Datacenter Networks
Wanchun Jiang, Haoyang Li 0006, Danfeng Shan, Fengyuan Ren, Jiawei Huang 0001, Jianxin Wang 0001 |
NSDI | 6 |
| 2026 | REAL: Emulating Control Plane at Simulator's Cost
Ze Xia, Hao Li 0011, Jinyu Fu, Yihan Dang, Danfeng Shan, Li Chen 0008, Peng Zhang 0011 |
NSDI | 6 |
| 2026 | Nüwa: A Generative Control Plane for AI Network SimulationabstractNetwork simulation plays a critical role in improving the efficiency of large-scale AI clusters for design validation, parameter tuning, and protocol development. However, high-fidelity network simulation becomes prohibitively slow at scale, especially when running large batches of experiments on topologies with tens or hundreds of thousands of accelerators. We observe that a key bottleneck comes from the control plane. Existing network simulators typically compute routes and install forwarding tables at initialization, which can consume hundreds of GB of memory before packet-event execution begins and limit overall simulation throughput. In this paper, we present Nüwa, which views routing as a compilation problem, it leverages the hierarchical and symmetric structure common in AI fabrics and compiles a declarative topology description together with routing policies into compact forwarding artifacts that are fast to generate and efficient to look up. Evaluations show that Nüwa can reduce simulation initialization time from hours to only 25 seconds for a 65,536-GPU cluster. For end-to-end simulation time, Nüwa takes only 20% of that required by existing approaches in a 40K+ GPU cluster, and Nüwa can scale to a 221,184-GPU cluster. Ran Shu 0001, Peng Zhang 0011, Danfeng Shan, Yongqiang Xiong |
SIGCOMM | 5 |
| 2026 | Fast and Accurate Software Traffic Shaping With Inter-Flow Batching
Danfeng Shan, Shihao Hu, Hao Li 0011, Yazhe Tang, Peng Zhang 0011, Wanchun Jiang, Fengyuan Ren |
IEEE Trans. Netw. | 1 |
| 2026 | Efficient Headroom Allocation With Two-Level Flow Control for Lossless Datacenter NetworksabstractIn datacenters, lossless network is very attractive as it can achieve ultra-low latency. In commodity Ethernet, lossless forwarding is achieved by hop-by-hop Priority-based Flow Control (PFC). To avoid buffer overflow, PFC-enabled switches need to reserve some buffer asheadroom, absorbing in-flight packets during the delay for backpressure messages to take effect. However, with the growing link speed in production networks, the buffer becomes increasingly insufficient, and the headroom can occupy a considerable fraction of buffer. As a result, the remaining buffer for absorbing normal traffic bursts is significantly squeezed, leading to frequent PFC messages that degrade the network performance. Worse yet, we find that the current static and queue-independent headroom allocation scheme is quite inefficient, resulting in significant buffer wastage. In light of this, we propose Dynamic and Shared Headroom allocation scheme (DSH), which dynamically allocates headroom to congested queues and enables sharing of allocated headroom among different queues. To achieve this, DSH first introduces port-level flow control, which performs flow control at the granularity of individual ports, guaranteeing lossless forwarding with a small fraction of per-port headroom. With this lossless guarantee, the switch is liberated for dynamic headroom adjustment. DSH dynamically allocates per-queue headroom based on the congestion status of each queue. Meanwhile, DSH preserves the queue-level flow control to protect the non-congested queues from being paused by congested queues, ensuring performance isolation on buffer sharing. Extensive experiments show that DSH can reduce the flow completion time by up to ~78.8%. Danfeng Shan, Jinchao Ma, Yunguang Li, Boxuan Hu, Tong Zhang 0018, Yazhe Tang, Hao Li 0011, Jinyu Wang 0002, Peng Zhang 0011 |
IEEE Trans. Netw. | 1 |
| 2025 | Unleashing the Power of Visual Foundation Models for Generalizable Semantic SegmentationabstractDeep learning models often suffer from performance degradation in unseen domains, posing a risk for safety-critical applications such as autonomous driving. To tackle this problem, recent studies have leveraged pre-trained Visual Foundation Models (VFMs) to enhance generalization. However, exsiting works mainly focus on designing intricate networks for VFMs, neglecting their inherent strong generalization potential. Moreover, these methods typically perform inference on low-resolution images. The loss of detail hinders accurate predictions in unseen domains, especially for small objects. In this paper, we argue that simply fine-tuning VFMs and leveraging high-resolution images unleash the power of VFMs for generalizable semantic segmentation. Therefore, we design a VFM-based segmentation network (VFMNet) that adapts VFMs to this task with minimal fine-tuning, preserving their generalizable knowledge. Then, to fully utilize high-resolution images, we train a Mask-guided Refinement Network (MGRNet) to refine VFMNet's predictions combining detailed image features. Furthermore, we adopt a two-stage coarse-to-fine inference approach. MGRNet is used to refine the low-confidence regions predicted by VFMNet to obtain fine-grained results. Extensive experiments demonstrate the effectiveness of our method, outperforming state-of-the-art methods by 3.3% on the average mIoU in synthetic-to-real domain generalization. Peiyuan Tang, Xiaodong Zhang 0036, Chunze Yang, Haoran Yuan, Jun Sun 0001, Danfeng Shan, Zijiang Yang 0006 |
AAAI | 6 |
| 2025 | Occamy: A Preemptive Buffer Management for On-chip Shared-memory SwitchesabstractToday's high-speed switches employ an on-chip shared packet buffer. The buffer is becoming increasingly insufficient as it cannot scale with the growing switching capacity. Nonetheless, the buffer needs to face highly intense bursts and meet stringent performance requirements for datacenter applications. This imposes rigorous demand on the Buffer Management (BM) scheme, which dynamically allocates the buffer across queues. However, the de facto BM scheme, designed over two decades ago, is ill-suited to meet the requirements of today's network. Danfeng Shan, Yunguang Li, Jinchao Ma, Xinyu Wen, Hao Li 0011, Wanchun Jiang, Nan Li 0047, Fengyuan Ren |
EuroSys | 1 |
| 2025 | Ditto: A Flexible Transmission Control Mechanism for Diverse User Demands and Network Conditions
Qingnan Wang, Jinnuo Du, Fangzhou Chen, Danfeng Shan |
ICA3PP (4) | 4 |
| 2025 | Bandwidth Prediction Within Playback-Time of Buffer Duration Based on Frame-Level Stable Segment for Live Streaming
Xianshi Su, Wanchun Jiang, Danfeng Shan |
WASA (2) | 4 |
| 2025 | Teaching to Fish Rather Than Giving a Fish: The Concentrator Method of Teaching Classic Congestion Control With Learning-Based moduleabstractNowadays, Congestion Control (CC) algorithms are expected to satisfy the diverse demands of applications running over diverse networks. To achieve this goal, the combinations, which are expected to inherit both the advantages of classic CC in terms of convergence, overhead, and explainability, and the advantages of learning-based CC on adapting to diverse networks and demands, become a hot topic. In this paper, we reveal the existing combination works are eithergiving a fishorteaching to fish. Based on the insight of their essential issues, we develop the Concentrator method ofteaching to fish. According to this method, we propose Seagull as a step further. Specifically, Seagull captures the network characteristics and application demands in a coarse-grained manner via an online learning module. Moreover, the online learning module guides the customization of the rate adjustment rules of the classic CC module for fine-grained system evolution. Replacing the assumption on networks by the captured characteristics, the classic CC module of Seagull can fulfill the specified application demands. Real-world experimental results show Seagull respectively outperforms Orca, PCC-Vivace, and CUBIC by$49.3\%,\ 30.4\%$, and 24.9% in terms of throughput over the Internet, and improves the video quality of experience (QoE) by$12.9\sim 33.5\%$compared to CUBIC over cellular links. Haoyang Li 0006, Wanchun Jiang, Jie Wang 0067, Jiawei Huang 0001, Danfeng Shan, Jianxin Wang 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | Improvement of Copa: Behaviors and Friendliness of Delay-Based Congestion Control AlgorithmabstractDelay-based congestion control has drawn a lot of attention in both academics and industry recently. Specifically, the Copa algorithm proposed in NSDI can achieve consistent high performance under various network environments and has already been deployed on Facebook. In this paper, we theoretically analyze Copa and reveal its large queuing delay and poor fairness issue under certain conditions. The root cause is that Copa fails to achieve its expected behaviors, i.e., clear the bottleneck buffer occupancy periodically. Moreover, we also reveal that the pathological competitive mode of Copa fails to guarantee friendliness. To address these issues, we propose Copa+, which enhances Copa with a parameter adaptation mechanism and an optimized competitive mode. Designed based on our theoretical analysis, Copa+ can adaptively clear the bottleneck buffer occupancy and become friendly to Cubic in the competitive mode. As a result, Copa+ inherits the advantages of Copa but achieves lower queuing delay and better fairness under different environments, as confirmed by real-world experiments and simulations. Specifically, Copa+ has the highest average throughput over different Internet links among different cloud nodes, compared to Cubic, BBR, PCC Vivace, Remy, and Indigo. Meanwhile, Copa+ has an 8.1% increase in throughput and similar low queuing delay compared to Copa. Moreover, Copa+ achieves 14.6% lower queuing delay and 2.4% higher throughput compared to Sprout over emulated cellular links. Wanchun Jiang, Haoyang Li 0006, Jia Wu 0002, Zheyuan Liu 0008, Jiawei Huang 0001, Danfeng Shan, Jianxin Wang 0001 |
IEEE/ACM Trans. Netw. | 6 |
| 2024 | Programming Network Stack for Physical Middleboxes and Virtualized Network FunctionsabstractMiddleboxes are becoming indispensable in modern networks. However, programming the network stack of middleboxes to support emerging transport protocols and flexible stack hierarchy is still a daunting task. To this end, we propose Rubik, a language that greatly facilitates the task of middlebox stack programming. Different from existing hand-written approaches, Rubik offers various high-level constructs for relieving the operators from dealing with massive native code, so that they can focus on specifying their processing intents. We show that using Rubik one can program the middlebox stack with minor effort, e.g., 250 lines of code for a complete TCP/IP stack, which is a reduction of 2 orders of magnitude compared to the hand-written versions. To maintain a high performance, we conduct extensive optimizations at the middle-and back-end of the compiler. Experiments show that the stacks generated by Rubik outperform the mature hand-written stacks by at least 30% in throughput. Hao Li 0011, Yihan Dang, Guangda Sun, Changhao Wu, Peng Zhang 0011, Danfeng Shan, Tian Pan 0001, Chengchen Hu |
IEEE/ACM Trans. Netw. | 6 |
| 2024 | Enforcing Fairness in the Traffic Policer Among Heterogeneous Congestion Control AlgorithmsabstractTraffic policing is widely used by ISPs to limit their customers’ traffic rates. It has long been believed that a well-tuned traffic policer offers a satisfactory performance for TCP. However, we find this belief breaks with the emergence of new congestion control (CC) algorithms: flows using new CC algorithms can easily occupy the majority of bandwidth, starving traditional TCP flows. We confirm this problem with experiments and reveal its root cause as follows. Without a buffer in traffic policers, congestion only causes packet losses, while new CC algorithms are loss-resilient. When being policed, they will not reduce the sending rate until an unacceptable loss ratio for TCP is reached, resulting in low throughput for competing TCP flows. Simply adding a buffer to the traffic policer improves fairness but incurs high latency. To this end, we propose FairPolicer, which can achieve fair bandwidth allocation without sacrificing latency. FairPolicer regards a token as a basic unit of bandwidth and fairly allocates tokens to active flows in a round-robin manner. To avoid bandwidth waste when flows come and go, FairPolicer puts all available tokens in a global bucket and maintains the amount of residual bucket space rather than the number of available tokens. To scale to massive concurrent flows, FairPolicer uses a Count-Min Sketch structure to maintain per-flow data with a small memory footprint. Testbed experiments show that FairPolicer can allocate bandwidth in a max-min fair manner and achieve much lower latency than other kinds of rate limiters. Danfeng Shan, Linbing Jiang, Peng Zhang 0011, Wanchun Jiang, Hao Li 0011, Yazhe Tang, Fengyuan Ren |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | Less is More: Dynamic and Shared Headroom Allocation in PFC-Enabled Datacenter NetworksabstractIn datacenters, lossless network is very attractive as it can achieve ultra-low latency. In commodity Ethernet, lossless forwarding is achieved by hop-by-hop Priority-based Flow Control (PFC). To avoid buffer overflow, PFC-enabled switches need to reserve some buffer as headroom, which is for absorbing in-flight packets during the delay for backpressure messages to take effect. However, with the growing link speed in production networks, the buffer becomes increasingly insufficient, and the headroom can occupy a considerable fraction of buffer. As a result, the remaining buffer for absorbing normal traffic bursts is significantly squeezed, leading to frequent PFC messages that degrade the network performance. However, the current static and queue-independent headroom allocation scheme is inherently inefficient in solving this problem. In light of this, we propose Dynamic and Shared Headroom allocation scheme (DSH), which dynamically allocates headroom to congested queues and enables the allocated headroom to be shared among different queues. By statistical multiplexing, DSH needs much less headroom to ensure lossless forwarding. Furthermore, DSH can be implemented on switching chips with moderate modifications. Extensive simulations show that DSH can absorb 4× more bursts without triggering PFC messages and reduce the flow completion time by up to ~31%. Danfeng Shan, Tong Zhang 0018, Yazhe Tang, Hao Li 0011, Peng Zhang 0011 |
ICDCS | 1 |
| 2023 | Reinforcement-Learning Based Preload Strategy for Short Video
Zhicheng Ren, Yongxin Shan, Wanchun Jiang, Yijing Shan, Danfeng Shan, Jianxin Wang 0001 |
ICIC (5) | 5 |
| 2023 | weBurst can be Harmless: Achieving Line-rate Software Traffic Shaping by Inter-flow BatchingabstractTraffic shaping is a common function at end hosts. Compared with hardware ones, software shapers are more flexible to be developed and deployed, and thus are very attractive. Nevertheless, software approaches are still unsatisfactory as they struggle to saturate 40Gbps and higher speed.While much effort has been made to reduce the intrinsic overhead of software traffic shaping, we find that it is the extrinsic overhead, such as PCIe communications and interrupts, that hinders shaping from achieving 40Gbps - 100Gbps speed. Batching is an effective way to amortize these overheads. However, blindly batching can degrade the network performance, as it introduces bursts into the network. Diving into the dilemma, we find that intra-flow burst is to blame for harming the network performance, while inter-flow burst, consisting of packets from different flows, can be naturally demultiplexed in the network.Based on the insight, we present FlowBundler, which can achieve efficient traffic shaping by inter-flow batching. Testbed experiments show that FlowBundler can achieve an accurate shaping of 98Gbps with a single CPU core, which is 2.6× better than state-of-the-art approaches. Large-scale simulations show that FlowBundler can batch packet transmissions without harming the network performance. Danfeng Shan, Shihao Hu, Wanchun Jiang, Hao Li 0011, Peng Zhang 0011, Yazhe Tang, Huanzhao Wang, Fengyuan Ren |
INFOCOM | 1 |
| 2023 | LemonNFV: Consolidating Heterogeneous Network Functions at Line Speed
Hao Li 0011, Yihan Dang, Guangda Sun, Guyue Liu, Danfeng Shan, Peng Zhang 0011 |
NSDI | 5 |
| 2022 | Copa+: Analysis and Improvement of the Delay-based Congestion Control Algorithm CopaabstractCopa is a delay-based congestion control algorithm proposed in NSDI recently. It can achieve consistent high performance under various network environments and has already been deployed in Facebook. In this paper, we theoretically analyze Copa and reveal its large queuing delay and poor fairness issue under certain conditions. The root cause is that Copa fails to clear the bottleneck buffer occupancy periodically as expected. Accordingly, Copa may get a wrong base RTT estimation and enter its competitive mode by mistake, leading to large delay and unfairness. To address these issues, we propose Copa+, which enhances Copa with a parameter adaptation mechanism and an optimized competitive mode entrance criterion. Designed based on our theoretical analysis, Copa+ can adaptively clear the bottleneck buffer occupancy for correct estimation of base RTT. Consequently, Copa+ inherits the advantages of Copa but achieves lower queuing delay and better fairness under different environments, as confirmed by the real-world experiments and simulations. Specifically, Copa+ has the highest throughput similar to Copa but 11.9% lower queuing delay over different Internet links among different cloud nodes, and achieves 39.4% lower queuing delay and 8.9% higher throughput compared to Sprout over emulated cellular links. Wanchun Jiang, Haoyang Li 0006, Zheyuan Liu 0008, Jia Wu 0002, Jiawei Huang 0001, Danfeng Shan, Jianxin Wang 0001 |
INFOCOM | 6 |
| 2022 | Demystifying and Mitigating TCP CappingabstractToday’s Internet user experience greatly depends on some user-perceived network metrics, such as throughput and latency. To improve these metrics, many Internet content providers build the content delivery network (CDN) to provide their services. Generally, CDNs adopt TCP as their transport protocol. A recent line of work improves TCP by proposing novel congestion control algorithms. However, we measure TCP performance in the production CDN and identify an interesting phenomenon termed TCP capping. When the flows experience TCP capping, the fixed-size receive window (rwnd) restricts these flows from fully utilizing network bandwidth. Through in-depth analysis, we demystify that the root cause of TCP capping is an inappropriate constraint on rwnd due to not considering the receiver’s processing capability. To mitigate it, this paper proposes a server-side scheme Apollo and a client-side scheme Artemis for Internet content providers and users, respectively. Apollo probes the receiver’s processing capability and assists the sender in packet sending. And Artemis adjusts the receive buffer in light of the receiver’s processing capability. In our evaluation, compared to vanilla TCP, TCP (w/ Apollo) and TCP (w/ Artemis) shorten flow completion time by up to 91.8% and 94.9%, respectively. Qingkai Meng 0001, Fengyuan Ren, Tong Zhang 0018, Danfeng Shan, Yajun Yang |
IWQoS | 4 |
| 2022 | Compiling Cross-Language Network Programs Into Hybrid Data PlaneabstractNetwork programming languages (NPLs) empower operators to program network data planes (NDPs) with unprecedented efficiency. Currently, various NPLs and NDPs coexist and no one can prevail over others in the short future. Such diversity is raising many problems including: (1) programs written with different NPLs can hardly interoperate in the same network, (2) most NPLs are bound to specific NDPs, hindering their independent evolution, and (3) compilation techniques cannot be readily reused, resulting in much wasteful work. These problems are mostly owing to the lack of modularity in the compilers, where the missing part is an intermediate representation (IR) for NPLs. To this end, we proposeNetwork Transaction Automaton (NTA), a highly-expressive and language-independent IR, and show it can express semantics of 7 mainstream NPLs. Then, we designCODER, a modular compiler based on NTA, which currently supports 2 NPLs and 3 NDPs. Experiments with real and synthetic programs show CODER can correctly compile those programs for real networks within moderate time. Hao Li 0011, Peng Zhang 0011, Guangda Sun, Wanyue Cao, Chengchen Hu, Danfeng Shan, Tian Pan 0001, Qiang Fu 0011 |
IEEE/ACM Trans. Netw. | 6 |
| 2021 | Towards the Fairness of Traffic PolicerabstractTraffic policing is widely used by ISPs to limit their customers' traffic rates. It has long been believed that a well-tuned traffic policer offers a satisfactory performance for TCP. However, we find this belief breaks with the emergence of new congestion control (CC) algorithms like BBR: flows using these new CC algorithms can easily occupy the majority of the bandwidth, starving traditional TCP flows. We confirm this problem with experiments and reveal its root cause as follows. Without buffer in traffic policers, congestion only causes packet losses, while new CC algorithms are loss-resilient, i.e. they adjust the sending rate based on other network feedback like delay. Thus, when being policed they will not reduce the sending rate until an unacceptable loss ratio for TCP is reached, resulting in low throughput for TCP. Simply adding buffer to the traffic policer improves fairness but incurs high latency. To this end, we propose FairPolicer, which can achieve fair bandwidth allocation without sacrificing latency. FairPolicer regards token as a basic unit of bandwidth and fairly allocates tokens to active flows in a round-robin manner. Testbed experiments show that FairPolicer can significantly improve the fairness and achieve much lower latency than other kinds of rate-limiters. Danfeng Shan, Peng Zhang 0011, Wanchun Jiang, Hao Li 0011, Fengyuan Ren |
INFOCOM | 1 |
| 2021 | Programming Network Stack for Middleboxes with Rubik
Hao Li 0011, Changhao Wu, Guangda Sun, Peng Zhang 0011, Danfeng Shan, Tian Pan 0001, Chengchen Hu |
NSDI | 5 |
| 2021 | RICH: Strategy-proof and efficient coflow scheduling in non-cooperative environments
Yazhe Tang, Danfeng Shan, Huanzhao Wang, Chengchen Hu |
J. Netw. Comput. Appl. | 3 |
| 2020 | An Intermediate Representation for Network Programming LanguagesabstractNetwork programming languages (NPLs) empower operators to program network data planes (NDPs) with unprecedented efficiency. Currently, various NPLs and NDPs coexist and no one can prevail over others in the short future. Such diversity is raising many problems including: (1) programs written with different languages can hardly interoperate in the same network, and (2) most NPLs are bound to specific NDPs, hindering their independent evolution. These problems are mostly owing to the lack of modularity in the compilers, where the missing part is an intermediate representation (IR) for NPLs. To this end, we propose Network Transaction Automaton (NTA), a highly-expressive and language-independent representation as the IR. We show that NTA can express semantics of 6 mainstream NPLs, and can be composed efficiently without any semantics loss. Hao Li 0011, Peng Zhang 0011, Guangda Sun, Chengchen Hu, Danfeng Shan, Tian Pan 0001, Qiang Fu 0011 |
APNet | 5 |
| 2020 | One Rein to Rule Them All: A Framework for Datacenter-to-User Congestion ControlabstractToday, considerable Internet traffic is sent from datacenter and heads for users. The network characteristics of connections served by servers in datacenters are usually diverse. As a result, a specific congestion control algorithm hardly accommodates the heterogeneity and performs well in various scenarios. In this work, we present Rein — a novel framework for Internet congestion control. With Rein, diverse congestion control algorithms can be assigned purposely to connections in one server to adapt to heterogeneity. We design and implement Rein in Linux, and the experiments validate that Rein is capable of smoothly switching among various candidate algorithms on the fly to achieve potential performance gain. Meanwhile, the overheads introduced by Rein are moderate and acceptable. Danfeng Shan, Xiaohui Luo, Tong Zhang 0018, Yajun Yang, Fengyuan Ren |
APNet | 2 |
| 2020 | A modular compiler for network programming languagesabstractNetwork programming languages (NPLs) empower operators to program network data planes (NDPs) with unprecedented efficiency. Currently, various NPLs and NDPs coexist and no one can prevail over others in the short future. Such diversity is raising many problems including: (1) programs written with different NPLs can hardly interoperate in the same network, (2) most NPLs are bound to specific NDPs, hindering their independent evolution, and (3) compilation techniques cannot be readily reused, resulting in much wasteful work. These problems are mostly owing to the lack of modularity in the compilers, where the missing part is an intermediate representation (IR) for NPLs. To this end, we propose Network Transaction Automaton (NTA), a highly-expressive and language-independent IR, and show it can express semantics of 7 mainstream NPLs. Then, we design CODER, a modular compiler based on NTA, which currently supports 2 NPLs and 3 NDPs. Experiments with real and synthetic network programs show CODER is efficient and scalable. Hao Li 0011, Peng Zhang 0011, Guangda Sun, Chengchen Hu, Danfeng Shan, Tian Pan 0001, Qiang Fu 0011 |
CoNEXT | 5 |
| 2020 | Observing and Mitigating Micro-Burst Traffic in Data Center NetworksabstractMicro-burst traffic is not uncommon in data centers. It can cause packet dropping, which may result in serious performance degradation (e.g., Incast problem). However, current approaches to mitigate micro-burst is usually ad-hoc and not based on a principled understanding of the underlying behaviors. On the other hand, traditional studies focus on traffic burstiness in a single flow, while micro-burst traffic in the data centers could occur with highly fan-in communication pattern, and its dynamic behavior is still unclear. To this end, in this paper, we re-examine the micro-burst traffic in typical data center scenarios. We find that the evolution of micro-burst is determined by both TCP's self-clocking mechanism and congestion control algorithm. Besides, dynamic behaviors of micro-burst under various scenarios can all be described by the time derivative of queue length evolution.Our observations also implicate that conventional solutions like absorbing and pacing are ineffective to mitigate micro-burst traffic.Instead, senders need to rapidly respond to some explicit signals of the queue buildup caused by the micro-burst traffic rather than independently and ineffectually pacing themselves in isolation. Inspired by the findings and insights from experimental observations, we propose Micro-burst-Aware Transport Control Protocol (MATCP), which leverages characteristic behaviors of micro-burst traffic derived from the time derivative of the queue occupancy. MATCP can suppress the sharp queue length increment by over 2x and reduce the tail query completion time by up to 84.4%. Danfeng Shan, Fengyuan Ren, Peng Cheng 0005, Ran Shu 0001, Chuanxiong Guo |
IEEE/ACM Trans. Netw. | 1 |
| 2018 | Estimating Short Connection Capacity on High Performance User Level Network StackabstractShort connections are generally used to transfer small-size messages, which contribute a large part of workload in modern applications. The maximum sustainable short connection rate, which is called short connection capacity, is an important index for admission control, Web QoS control, and energy saving. A capacity estimation mechanism aims to find the workload just saturating the server, and it relies on both workload information and system information. Past researches point out that kernel space network stack becomes the bottleneck when a huge number of concurrent short connections coexist. On the other hand, high performance user level network stacks have been proved to eliminate such bottleneck, thus become a hot research topic in both academia and industry. However, they also bring challenges for estimating short connection capacity, making traditional methods ineffective. Therefore, it is important to find a new method to estimate short connection capacity on high performance user level network stacks. In this paper, we prove that the effective CPU utilization is an adaptive index to different workload patterns and application complexities, which can reflect the server state. Then we design and implement an online capacity estimator on the Seastar platform. We conduct experiments to verify the effectiveness of our online capacity estimator. The results show that our estimator can actually estimate the capacity online. When the server is near saturated, the 90th percentile relative estimating error is no more than 9.18%. Furthermore, our capacity estimator only introduces no more than 1.38% of capacity loss in our experiments. Jing Xie 0005, Wenxue Cheng, Tong Zhang 0018, Danfeng Shan, Fengyuan Ren |
ICCCN | 4 |
| 2018 | Micro-Burst in Data Centers: Observations, Analysis, and MitigationsabstractMicro-burst traffic is not uncommon in data centers. It can cause packet dropping, which results in serious performance degradation (e.g., Incast problem). However, current solutions that attempt to suppress micro-burst traffic are extrinsic and ad hoc, since they lack the comprehensive and essential understanding of micro-burst's root cause and dynamic behavior. On the other hand, traditional studies focus on traffic burstiness in a single flow, while in data centers micro-burst traffic could occur with highly fan-in communication pattern, and its dynamic behavior is still unclear. To this end, in this paper, we re-examine the microburst traffic in typical data center scenarios. We find that evolution of micro-burst is determined by both TCP's self-clocking mechanism and bottleneck link. Besides, dynamic behaviors of micro-burst under various scenarios can all be described by the slope of queue length evolution. Our observations also implicate that conventional solutions like absorbing and pacing are ineffective to mitigate micro-burst traffic. Instead, senders need to slow down as soon as possible. Inspired by the findings and insights from experimental observations, we propose S-ECN policy, which is an ECN marking policy leveraging the slope of queue length evolution. Transport protocols utilizing S-ECN policy can suppress the sharp queue length increment by over 2×, and reduce the average query completion time by ~12-27%. Danfeng Shan, Fengyuan Ren, Peng Cheng 0005, Ran Shu 0001, Chuanxiong Guo |
ICNP | 1 |
| 2018 | ECN Marking With Micro-Burst Traffic: Problem, Analysis, and Improvement
Danfeng Shan, Fengyuan Ren |
IEEE/ACM Trans. Netw. | 1 |
| 2018 | MPTCP Tunnel: An Architecture for Aggregating Bandwidth of Heterogeneous Access NetworksabstractFixed and cellular networks are two typical access networks provided by operators. Fixed access network is widely employed; nevertheless, its bandwidth is sometimes not sufficient enough to meet user bandwidth requirements. Meanwhile, cellular access network owns unique advantages of wider coverage, faster increasing link speed, more flexible deployment, and so forth. Therefore, it is attractive for operators to mitigate the bandwidth shortage by bundling these two. Actually, there have been existing schemes proposed to aggregate the bandwidth of two access networks, whereas they all have their own problems, like packet reordering or extra latency overhead. To address this problem, we design new architecture, MPTCP Tunnel, to aggregate the bandwidth of multiple heterogeneous access networks from the perspective of operators. MPTCP Tunnel uses MPTCP, which solves the reordering problem essentially, to bundle multiple access networks. Besides, MPTCP Tunnel sets up only one MPTCP connection at play which adapts itself to multiple traffic types and TCP flows. Furthermore, MPTCP Tunnel forwards intact IP packets through access networks, maintaining the end‐to‐end TCP semantics. Experimental results manifest that MPTCP Tunnel can efficiently aggregate the bandwidth of multiple access networks and is more adaptable to the increasing heterogeneity of access networks than existing mechanisms. Danfeng Shan, Ran Shu 0001, Tong Zhang 0018 |
Wirel. Commun. Mob. Comput. | 2 |
| 2017 | Improving ECN marking scheme with micro-burst traffic in data center networksabstractIn data centers, micro-burst is a common traffic pattern. The packet dropping caused by it usually leads to serious performance degradations. Therefore, much attention has been paid to avoiding buffer overflow caused by micro-burst traffic. In particular, ECN is widely used in data centers to keep persistent queue occupancy low, so that enough buffer space can be available as headroom to absorb micro-burst traffic. However, we find that instantaneous-queue-length-based ECN may cause problems in another direction - buffer underflow. Specifically, current ECN marking scheme in data centers is easy to trigger spurious congestion signals, which may result in overreaction of senders and queue length oscillations in switches. Since ECN threshold is low, the buffer may underflow and link capacity is not fully used. In this paper, we reveal this problem by experiments and simulations. Besides, we theoretically deduce the amplitude of queue length oscillations. The analysis result shows that overreaction of senders is caused by ECN mis-marking. Therefore, we propose Combined Enqueue and Dequeue Marking (CEDM), which can mark packets more accurately. Through simulations, we show that CEDM can greatly reduce throughput loss and improve flow completion time. Danfeng Shan, Fengyuan Ren |
INFOCOM | 1 |
| 2017 | XpressEth: Concise and efficient converged real-time EthernetabstractOwing to Ethernet's low cost, high bandwidth and architecture openness, much attention has been paid to develop converged Ethernet to support both time-critical services and conventional communication services on a unified network infrastructure. The greatest challenge here is providing low and deterministic latency for time-critical packets. Recently, the IEEE time sensitive networking task group is launched to address it. However, their framework is complex and unsuitable for commodity switch architecture. In this paper, we propose a concise and efficient converged real-time Ethernet framework called XpressEth, which leverages Dual Preemption mechanism to minimize the delay of time-critical packets, and employs a lightweight Slot Assignment Scheduler to minimize the conflicts among time-critical packets at sources. XpressEth cuts off great burden from both forwarding and scheduling. The simulation results verify that XpressEth can provide ultra-low and deterministic latency for time-critical packets (1.024μ s per hop and zero jitter in 1Gbps network), which is 13× better than time sensitive networking solution, and the side-effect on conventional communication traffic is negligible. Kun Qian 0017, Fengyuan Ren, Danfeng Shan, Wenxue Cheng, Bo Wang 0066 |
IWQoS | 3 |
| 2017 | Analyzing and Enhancing Dynamic Threshold Policy of Data Center SwitchesabstractToday's data center switches usually employ on-chip shared memory; buffer management policy in them is essential to ensure fair sharing of memory among all ports. Among various polices, Dynamic Threshold (DT) policy is widely used by switch vendors. Meanwhile, in data centers, distributed applications such as MapReduce often introduce micro-burst traffic into network and the packet dropping caused by micro-burst usually leads to serious performance degradation. When micro-burst traffic arrives at switches, DT is unable to fully utilize the buffer to absorb it. Therefore, in this paper, we theoretically deduce the sufficient conditions for packet dropping caused by micro-burst traffic, and quantitatively estimate the free buffer size when packets are dropped. The results show that the free buffer size can be very large when the number of overloaded ports is small. What's worse, to ensure fair sharing of memory among output ports, packets from micro-burst traffic may be dropped even when the traffic size is much smaller than the buffer size. In light of these results, we propose the Enhanced Dynamic Threshold (EDT) policy, which can alleviate packet dropping caused by micro-burst traffic through fully utilizing the switch buffer and temporarily relaxing the fairness constraint. The simulation results show that EDT can absorb more micro-burst traffic than DT. Danfeng Shan, Wanchun Jiang, Fengyuan Ren |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | Absorbing micro-burst traffic by enhancing dynamic threshold policy of data center switchesabstractIn data center networks, micro-burst is a common traffic pattern and the packet dropping caused by it usually leads to serious performance degradation. Meanwhile, most of the current commodity switches employ on-chip shared memory, and the buffer management policies of them ensure fair sharing of memory among all ports. Among various polices, Dynamic Threshold (DT) is widely used by switch vendors. However, because DT needs to reserve a fraction of switch buffer, there is free buffer space while packets from micro-burst traffic are dropped. In this paper, we theoretically deduce the sufficient conditions for packet dropping caused by micro-burst traffic, and estimate the corresponding free buffer size. The results show that the free buffer size is very large when the number of overloaded ports is small. What's worse, to ensure fair sharing of memory among output ports, packets from micro-burst traffic may be dropped even when the traffic size is much smaller than the buffer size. In light of these results, we propose Enhanced Dynamic Threshold (EDT) policy, which can alleviate packet dropping caused by micro-burst traffic through fully utilizing the switch buffer and temporarily relaxing the fairness constraint. The simulation results show that EDT can absorb more micro-burst traffic than DT. Danfeng Shan, Wanchun Jiang, Fengyuan Ren |
INFOCOM | 1 |
| 2013 | Inter-Swarm Content Distribution Among Private BitTorrent NetworksabstractPrivate BitTorrent (PT) is a new trend in Peer-to-Peer file sharing system, which provides high incentives for its users to seed after download by maintaining an upload-to-download ratio in the tracker for each registered community member. From the data we collected from six active PT sites, we discover that the population of both users and contents in any single PT site is much less than the public BitTorrent, and the intersection of content sets in different PTs is quite small. Based on this observation, we propose a content sharing/distribution framework among PTs (named CrossPT), as well as its sharing mechanism. In addition, we investigate the sharing strategy of the PT participants in CrossPT using game theory and the fetch strategy by modeling the scenario to a Neighbor Selection Problem (NSP). We prove NSP to be NP-complete and propose a heuristic algorithm to solve it. The evaluations with the input of crawled data from six PT sites demonstrate the efficiency of our mechanism. The content sizes of the six PT sites can be increased by 113.95%-438.46% with CrossPT. Also, the content distribution process can be done in less than one second, excluding the delivery time of the content itself. Chengchen Hu, Danfeng Shan, Yu Cheng 0003, Tao Qin 0002 |
IEEE J. Sel. Areas Commun. | 2 |