EDBT 2026 Demo / reviewers in the wild / expert
Qingkai Meng 0001
dblp:54/2449-1
· DBLP profile ↗
31ranked-venue papers
7as first author
28since 2021 · last 2026
0000-0001-5394-8450ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 25 · 7 first-author · 24 since 2021Systems, architecture and hardware · 5 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking MoE Routing for Commodity GPUsabstractCommodity GPUs make large-model serving economically attractive, but they are typically connected via PCIe rather than high-bandwidth GPU fabrics. This is a poor fit for expert-parallel Mixture-of-Experts (MoE) inference, where each MoE layer routes tokens to selected experts via all-to-all communication, and PCIe crossings can become a major latency bottleneck. Existing serving systems and communication libraries optimize this traffic but treat router-selected expert assignments as fixed. We propose PCIe-MoE, an inference-time mechanism that decides when a low-confidence remote expert is worth a costly PCIe hop, and when a nearby, score-similar expert can be substituted instead. PCIe-MoE combines topology profiling, replaceability-aware placement, risk-aware remapping, and quality guards to reshape token-to-expert routing without retraining. Meng Li 0010, Siyuan Tong, Qingkai Meng 0001, Xiaoliang Wang 0001, Haipeng Dai 0001 |
APNet | 4 |
| 2026 | Fine-grained and Non-intrusive LLM Training Monitoring via Microsecond-level Traffic MeasurementabstractLarge language model (LLM) training is prone to anomalies due to its long duration and large scale, which can lead to significant performance degradation or even training crashes. Due to the synchronization nature of LLM training, anomalies exhibit the cascading effect, making their diagnosis challenging. Existing approaches rely on collecting communication operator information via code instrumentation, which yields only coarse-grained monitoring data and requires modifications to training code or communication libraries. We propose Pulse, a fine-grained, non-intrusive, and easy-to-deploy monitoring system. Our key idea is to enable fine-grained monitoring via traffic measurement. Pulse conducts microsecond-level RDMA traffic measurement on NICs, and transforms flow-level measurements into communication operator measurements, thereby enabling fine-grained and non-intrusive monitoring. We deploy Pulse on a testbed with 64 H200 GPUs and evaluate its anomaly localization capability under common failure scenarios. Pulse achieves machine-level localization in 10 out of 12 scenarios, while existing methods succeed in only 4 and even misdiagnose 2 of the remaining scenarios. Additionally, Pulse achieves over 90% precision and 100% recall, supports up to 2000 concurrent RDMA flow measurements per NIC, and imposes negligible overhead on training performance, making it a practical solution for real-world LLM training environments. Yibo Xiao, Haifeng Sun 0004, Qingkai Meng 0001, Jiong Duan, Xiaohe Hu, Rong Gu 0001, Guihai Chen, Chen Tian 0001 |
ASPLOS (2) | 4 |
| 2026 | STAR: Decode-Phase Rescheduling for LLM InferenceabstractLarge Language Model (LLM) inference has emerged as a fundamental paradigm, however, variations in output length cause severe workload imbalance in the decode phase, particularly for long-output reasoning tasks. Existing systems, such as PD disaggregation architectures, rely on static prefill-to-decode scheduling, which often results in SLO violations and OOM failures under evolving decode workloads. In this paper, we propose STAR, a decode rescheduling system powered by length prediction to anticipate future workloads. Our core contributions include: (1) A lightweight and continuous LLM-native prediction method that leverages LLM hidden state to model remaining generation length with high precision (reducing MAE by 49.42%) and low overhead (cutting predictor parameters by 93.28%); (2) A rescheduling solution in decode phase with a dynamic balancing mechanism that integrates current and predicted workloads, reducing P99 TPOT by 75.1% and achieving 2.63 × higher goodput. Zhibin Wang 0002, Zetao Hong, Xue Li 0024, Qingkai Meng 0001, Qing Wang 0031, Chengying Huan, Rong Gu 0001, Sheng Zhong 0002, Chen Tian 0001 |
HPDC | 6 |
| 2026 | PRO: Deterministic Load Balancing for AI Datacenter Network
Qingkai Meng 0001, Xiangyu Han, Shangguang Wang |
IWQoS | 2 |
| 2026 | Anytest: Localizing the Root Cause of Hardware Transport Performance Anomalies
Zhaochen Zhang, Sheng Cheng 0002, Feiyang Xue, Chang Liu 0001, Boliang Liu, Rui Li 0020, Li Wang 0110, Peirui Cao, Qingkai Meng 0001, Guihai Chen, Shuguang Cheng, Yongqing Xi, Binzhang Fu, Dennis Cai, Chen Tian 0001 |
SIGCOMM | 13 |
| 2025 | ScalaTap: Scalable Outbound Rate Limiting in Public Cloud
Zhongjie Chen, Yingchen Fan, Kun Qian 0017, Qingkai Meng 0001, Ran Shu 0001, Bo Wang 0066, Wei Li 0262, Fengyuan Ren |
INFOCOM | 4 |
| 2025 | eTran: Extensible Kernel Transport with eBPF
Zhongjie Chen, Qingkai Meng 0001, ChonLam Lao, Fengyuan Ren, Minlan Yu, Yang Zhou 0008 |
NSDI | 2 |
| 2025 | Astral: A Datacenter Infrastructure for Large Language Model Training at ScaleabstractThe flourishing of Large Language Models (LLMs) calls for increasingly ultra-scale training. In this paper, we share our experience in designing, deploying, and operating our novel Astral datacenter infrastructure, along with operational lessons and evolutionary insights gained from its production use. Astral has three important innovations: (i) a same-rail interconnection network architecture on tier-2, which enables the scaling of LLM training. To physically deploy this high-density infrastructure, we introduce a distributed high-voltage direct current power system and a new air-liquid integrated cooling system. (ii) a full-stack monitoring system featuring cross-host and hierarchical logging correlation, which diagnoses failures at scale and precisely localizes root causes. (iii) an operator-granular forecasting component Seer that efficiently generates operator execution timelines with acceptable accuracy, aiding in fault diagnosis, model tuning, and network architecture upgrading. Astral infrastructure has been gradually deployed over 18 months, supporting LLM training and inference for multiple customers. Qingkai Meng 0001, Zhenhui Zhang, ChonLam Lao, Chengyuan Huang, Baojia Li 0002, Weizhen Dang, Zitong Lin, Yuanyuan Gong, Chunzhi He, Xiaoyuan Hu, Yinben Xia, Xiang Li 0223, Zekun He, Yachen Wang, Xianneng Zou, Kun Yang 0001, Gianni Antichi, Guihai Chen, Chen Tian 0001 |
SIGCOMM | 1 |
| 2025 | Incentivizing Fresh Information Acquisition via Age-Based RewardabstractMany Internet platforms are information-oriented and crowd-based. They collect fresh information of various points of interest (PoIs) relying on users who happen to be nearby the PoIs. The platform will offer rewards to incentivize users and compensate their costs incurred from information acquisition. In practice, a user’s cost is his/her private information, thus both the user cost and its distribution are hidden to the platform, making it challenging to determine the optimal rewarding decision. In this paper, we investigate how the platform dynamically rewards the users, aiming to jointly reduce the age of information (AoI) and the operational expenditure (OpEx). Due to the hidden cost distribution, this is an online non-convex learning problem with bandit feedback. To overcome the challenge, we first design an age-based reward scheme, which decouples the OpEx from the unknown cost distribution and enables the platform to accurately control its OpEx. We then take advantage of the age-based reward scheme and propose an exponentially discretizing and learning (EDAL) policy for platform operation. We prove that the EDAL policy performs asymptotically as well as the optimal decision (derived from the cost distribution). Simulation results show that the age-based reward scheme protects the platform’s OpEx from the influence of the user crowd characteristics, and also verify the asymptotic optimality of the EDAL policy. Zhiyuan Wang 0004, Qingkai Meng 0001, Shan Zhang 0001, Hongbin Luo |
IEEE Trans. Netw. | 2 |
| 2025 | Performance Evaluation for Latency-Stability Tradeoff in Satellite-Terrestrial Integrated NetworksabstractLow-Earth-Orbit (LEO) satellite constellation is promising to extend broadband Internet services to where the terrestrial networks cannot reach, forming a satellite-terrestrial integrated network (STIN). The tradeoff between communication latency and topology stability is the core of constellation design and STIN operation. In general, low-altitude orbits reduce the propagation delay of ground-satellite links (GSLs), but result in frequent handover for GSLs. Frequent GSL handover weakens the topology stability, which is often measured by satellite pass duration and communication session duration. In this paper, we investigate the latency-stability tradeoff in STIN, and propose a stochastic model to analyze and evaluate GSL latency and pass/session duration. We analytically derive the distributions and the expectations of these metrics in closed forms, and verify the theoretic results based on empirical distributions (obtained by trajectory simulation). Compared to previous studies, our results exhibit more concise expressions and arguments. We find that the above closed-form expressions take two key arguments, i.e., the orbit period and the maximal central angle. The expected pass/session duration is directly proportional to the product of the two arguments, and the coefficients can be obtained explicitly (e.g., 0.25 for pass duration). We also demonstrate how to utilize the analytic results to draw deep insights on constellation design, aiming to reduce the GSL latency and keep the topology stable. We believe that our results in this paper could empower the operators and researchers to obtain a rapid understanding on the latency-stability tradeoff of STIN without resorting to time-consuming simulations. Zhiyuan Wang 0004, Qingkai Meng 0001, Shan Zhang 0001, Hongbin Luo |
IEEE Trans. Netw. | 2 |
| 2025 | ACK-Driven Congestion Control for Lossless EthernetabstractCongestion control is a key enabler for lossless Ethernet at scale. In this paper, we revisit this classic topic from a new perspective, i.e., understanding and exploiting the intrinsic properties of the underlying lossless network. We experimentally and analytically find that the intrinsic properties of lossless networks, such as packet conservation, can indeed provide valuable implications in estimating pipe capacity and the precise number of excessive packets. Besides, we derive principles on how to treat congested flows and victim flows individually to handle HoL blocking efficiently. Then, we propose ACK-driven congestion control (ACC) for lossless Ethernet, which simply resorts to the knowledge of ACK time series (supports ACK coalescing) to exert a temporary halt to exactly drain out excessive packets of congested flows and then match its rate to pipe capacity. Testbed and large-scale simulations demonstrate that ACC ameliorates fundamental issues in lossless Ethernet (e.g., congestion spreading, HoL blocking, and deadlock) and achieves excellent low latency and high throughput performance. For instance, compared with existing schemes, ACC improves the average and 99th percentile FCT performance of small flows by$1.3\sim 3.3\times $and$1.4\sim 11.5\times $, respectively. Qingkai Meng 0001, Chaolei Hu, Shangguang Wang, Fengyuan Ren |
IEEE Trans. Netw. | 2 |
| 2024 | OpenSN: An Open Source Library for Emulating LEO Satellite NetworksabstractLow-earth-orbit (LEO) satellite constellations (e.g., Starlink) are becoming the necessary component of future Internet. There have been increasing studies on LEO satellite networking. It is a crucial problem how to evaluate these studies in a systematic and reproducible manner. In this paper, we present OpenSN, i.e., an open-source library for emulating large-scale satellite network (SN). Different from Mininet-based SN emulators (e.g., LeoEM), OpenSN adopts container-based virtualization, thus allows for running distributed routing software on each node, and can achieve horizontal scalability via flexible multi-machine extension. Compared to other container-based SN emulators (e.g., StarryNet), OpenSN streamlines the interaction with Docker command line interface and significantly reduces unnecessary operation of creating virtual links. These modifications improve emulation efficiency and vertical scalability on a single machine. Furthermore, OpenSN separates user-defined configuration from container network management via a key-value database that records the necessary information of SN emulation. Such a separation architecture enhances the function extensibility. To sum up, OpenSN exhibits advantages in efficiency, scalability, and extensibility, thus is a valuable open-source library that empowers research on satellite networking. Experimental results show that OpenSN can construct mage-constellations 5X-10X faster than StarryNet, and update link state 2X-4X faster than Mininet. Wenhao Lu, Zhiyuan Wang 0004, Shan Zhang 0001, Qingkai Meng 0001, Hongbin Luo |
APNet | 4 |
| 2024 | Adaptive Inter-Domain Content Retrieval in Satellite-Terrestrial Integrated NetworksabstractThe low-earth-orbit (LEO) satellite constellation is becoming the large-scale autonomous system (AS) connected to other terrestrial ASes, forming satellite-terrestrial integrated network (STIN). To reduce the latency of inter-domain content delivery, it is necessary to adopt an efficient inter-domain routing architecture in STIN. Previous studies improve the classic Border Gateway Protocol (BGP) and propose BGP-S for STIN, which proactively probes each inter-domain path. However, BGP-S suffers from a tremendous control overhead when the LEO constellation is large. In this paper, we propose a path-identified content retrieval (PCR) architecture for STIN. PCR is an information-centric architecture, and explicitly identifies the inter-domain path between neighbor ASes. Under PCR architecture, users can retrieve their desired contents by specifying the inter-domain path in the request packet. The corresponding data packets will be forwarded according to the reversed inter-domain path. This way, users could unleash the path diversity in STIN by dynamically selecting inter-domain paths according to various learning-based routing algorithms. That is, PCR achieves adaptive inter-domain content retrieval without incurring additional control overhead. We conduct extensive packet-level simulation on OMNeT++. Results show that PCR (with Exp3-SIX algorithm) outperforms state-of-the-art BGP-S in terms of content retrieval latency by 22%. Lixin Zeng, Zhiyuan Wang 0004, Shan Zhang 0001, Qingkai Meng 0001, Hongbin Luo |
GLOBECOM | 4 |
| 2024 | DockRDMA: Hybrid RDMA Virtualization for Containerized CloudsabstractContainers have become the de facto choice for major cloud services. Meanwhile, with demands for extremely high performance, data centers have widely adopted RDMA for their online services. RDMA virtualization is the critical technology that enables RDMA for containers. Hybrid RDMA virtualization leverages the software flexibility in the control path, and keeps the native performance in the data path. Thus, it is the best choice for RDMA virtualization. State-of-the-art hybrid RDMA virtualization cannot address containerspecific problems. This paper proposes DockRDMA, the first hybrid RDMA virtualization solution for containerized clouds. DockRDMA develops several mechanisms, including embedding physical addresses in virtual ones to provide efficient address translation, hybrid network policy enforcement at scale, a general virtual RDMA NIC initialization method to be compatible with all container platforms, and namespace checking to protect the RDMA NIC instances. Evaluation results show that DockRDMA provides bare-metal RDMA performance in the data path, and almost native communication setup time in the control path. Compared with the state-of-the-art hybrid virtualization technology, DockRDMA reduces Hadoop job completion time by 6%. It offers seamless integration with existing container platforms, protects critical information of RDMA NIC instances, and exhibits excellent scalability to meet diverse network policies required by different containers. Ran Shu 0001, Zhongjie Chen, Xiaohui Luo, Bo Wang 0066, Qingkai Meng 0001, Fengyuan Ren |
ICNP | 7 |
| 2024 | BCC: Re-architecting Congestion Control in DCNsabstractThe nature of datacenter traffic is a high volume of bursty tiny flows and standing long flows, which forms the coexistence of transient and persistent congestion. Traditional congestion control (CC) algorithms have inherent limitations in reconciling fast response and high efficiency towards transients with stability and fairness during persistence. In this paper, we provide an insight that re-architects CC with two control laws, tailored to transient and persistent concerns, respectively. Armed with this key insight, we propose bimodal congestion control (BCC), which is founded on two core ideas: (i) Quaternary network state detection, which further distinguishes transient and persistent states in switches, and (ii) Bimodal control law, which is manifested as the transient controller and persistent controller at sources. The transient controller employs a precise control paradigm that pauses flows to drain backlogged packets and ramps down/up flow rates to bottleneck bandwidth directly, striving for high efficiency. The persistent controller grounds itself in traditional CC algorithms, inheriting stability and fairness. We implement BCC in the Linux kernel and P4-programmable switch. In our evaluation, compared to DCQCN, HPCC, PowerTCP, and Swift, BCC reduces flow completion times by 14% ~ 99%. Qingkai Meng 0001, Shan Zhang 0001, Zhiyuan Wang 0004, Tao Tong, Chaolei Hu, Hongbin Luo, Fengyuan Ren |
INFOCOM | 1 |
| 2024 | Explicit Dropping Notification in Data CentersabstractDatacenter applications increasingly demand microsecond-scale latency and tight tail latency. Despite recent advances in datacenter transport protocols, we notice that the timeout caused by packet loss is the killer of microsecond-scale latency. Moreover, refining the RTO setting is impractical due to the significant fluctuations in RTT. In this paper, we propose explicit dropping notification (EDN) to avoid timeouts. EDN rekindles ICMP Source Quench, where the switch notifies the source of precise packet loss information. Then the source can rapidly pinpoint dropped packets for fast retransmission instead of waiting for timeouts. More importantly, fast retransmission does not mean immediate retransmission which is prone to aggravate congestion and deteriorate latency. In light of this, we suggest finessing the timing and sending rate of retransmission. Specifically, as a reward of the paradigm shift to explicit notification, the source can pause for the queue draining time piggybacked on EDN messages and estimate connection capacity to figure out a proper sending rate, thus avoiding congestion aggravation. We implement EDN on the P4-programmable switching ASIC and Linux kernel. Evaluations show that, compared with state-of-the-art loss recovery schemes, EDN reduces the latency by up to 4.1× on average and 3.6× at the 99th-percentile. Qingkai Meng 0001, Chaolei Hu, Bo Wang 0066, Fengyuan Ren |
INFOCOM | 1 |
| 2024 | Revisiting Congestion Control for Lossless Ethernet
Qingkai Meng 0001, Chaolei Hu, Fengyuan Ren |
NSDI | 2 |
| 2024 | Load-Aware Hierarchical Information-Centric Routing for Large-Scale LEO Satellite NetworksabstractThe emerging large-scale low earth orbit (LEO) constellation is expected to provide global Internet services. However, large-scale satellite networks meet the challenges of topology dynamics and continuous traffic variation. This paper proposes LoHi, a Load-aware Hierarchical Information-centric (LoHi) Routing Protocol based on constellation Partitioning and logic path identifier (PID), to address the above challenges. Specifically, LoHi divides a constellation into satellite groups and only keeps track of the inter-satellite link (ISL) state within each group instead of the entire constellation, stabilizing global routing. Moreover, LoHi adopts the PID to denote the logic connectivity of adjacent satellite groups, which corresponds to multiple physical ISLs. Accordingly, the hierarchical multipath routing is calculated based on the inner-group and inter-group connectivity information. Furthermore, LoHi adjusts forwarding decisions according to the real-time traffic load on the ISL and group-level inter-group paths to reduce packet drop and queuing delay in a hierarchical pattern. Packet-level experiment results show that LoHi achieves a higher packet delivery ratio than the state-of-the-art mechanisms (up to 132.12%). Zhiyuan Wang 0004, Shan Zhang 0001, Qingkai Meng 0001, Hongbin Luo |
WCNC | 4 |
| 2024 | Switch-Assistant Loss Recovery for RDMA Transport ControlabstractRoCEv2 (RDMA over Converged Ethernet version 2) is the canonical method for deploying RDMA in Ethernet-based datacenters. Traditionally, RoCEv2 runs over the lossless network which is in turn achieved by enabling Priority Flow Control (PFC) within the network. However, as the scale of the datacenter increases, PFC’s side effects, such as head-of-line blocking, congestion spreading, and pause frame storm, are amplified. Datacenter operators can no longer tolerate these problems. In hence, they are seeking PFC alternatives for RDMA networks. Rather than aiming at the lossless RDMA network, we instead handle packet loss effectively to support RDMA over Ethernet. In this paper, we propose Switch-assistant Loss Recovery (SLR), a switch building block to enhance RoCEv2’s loss recovery. Specifically, SLR-enabled switches send loss notifications to request fast retransmissions. To cooperate with go-back-N retransmission, SLR generates loss notifications only when expected packets (i.e., in-order packets expected by receivers) are dropped and then filters out unexpected packets, which can avoid timeouts and prevent exacerbating congestion. Further, we adapt SLR to multi-bottleneck scenarios by inferring expected packets among multiple switch views. We implement SLR prototypes on commodity programmable switches. Evaluations show that SLR reduces the 99.9th-percentile FCT slowdown by up to 21.6$\times$compared to PFC and other state-of-the-arts. Qingkai Meng 0001, Shan Zhang 0001, Zhiyuan Wang 0004, Tong Zhang 0018, Hongbin Luo, Fengyuan Ren |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | Enabling Byzantine Fault Tolerance in Access Authentication for Mega-ConstellationsabstractLow-Earth-Orbit (LEO) satellite constellations are becoming the necessary infrastructure in the future. However, the secure operation of LEO constellations is faced with severe risks. Specifically, LEO satellites are constantly orbiting and their channel interfaces are open. The adversary in hostile regions can leverage the global footprint to inject malicious traffic via access satellites. That is, LEO satellites are susceptible to physical and cyber attacks. Therefore, access authentication regarding terrestrial users (TUs) is crucial to ensure the secure operation of LEO constellations. The traditional on-orbit authentication frameworks usually presume that satellites are reliable and mutually trusted, thus one could rely on access satellites to perform authentication. In practice, however, physical and cyber attacks could bring down the satellites (causing fail-stop fault) or even hijack the satellites (causing Byzantine fault). This fact requires that the access authentication framework installed on LEO constellations should be fault-tolerant. In this paper, we aim to achieve Byzantine fault tolerance in access authentication for LEO satellite networks by properly integrating PBFT consensus protocol with traditional on-orbit authentication. Based on the topology characteristics of LEO constellations, we analytically derive the consensus probability, authentication accuracy, and communication overhead under PBFT-based authentication. To reduce the communication overhead, we propose to partition the constellation into multiple consensus groups, and devise a hierarchical PBFT (HPBFT) protocol. Simulation results based on Starlink Shell-I constellation indicate that HPBFT-based authentication could reduce the communication overhead (by an order of magnitude) and maintain almost the same authentication accuracy compared to PBFT-based authentication. Zhiyuan Wang 0004, Shan Zhang 0001, Qingkai Meng 0001, Hongbin Luo |
IEEE/ACM Trans. Netw. | 4 |
| 2023 | The Power of Age-based Reward in Fresh Information AcquisitionabstractMany Internet platforms collect fresh information of various points of interest (PoIs) relying on users who happen to be nearby the PoIs. The platform will offer reward to incentivize users and compensate their costs incurred from information acquisition. In practice, the user cost (and its distribution) is hidden to the platform, thus it is challenging to determine the optimal reward. In this paper, we investigate how the platform dynamically rewards the users, aiming to jointly reduce the age of information (AoI) and the operational expenditure (OpEx). Due to the hidden cost distribution, this is an online non-convex learning problem with partial feedback. To overcome the challenge, we first design an age-based rewarding scheme, which decouples the OpEx from the unknown cost distribution and enables the platform to accurately control its OpEx. We then take advantage of the age-based rewarding scheme and propose an exponentially discretizing and learning (EDAL) policy for platform operation. We prove that the EDAL policy performs asymptotically as well as the optimal decision (derived based on the cost distribution). Simulation results show that the age-based rewarding scheme protects the platform’s OpEx from the influence of the user characteristics, and verify the asymptotic optimality of the EDAL policy. Zhiyuan Wang 0004, Qingkai Meng 0001, Shan Zhang 0001, Hongbin Luo |
INFOCOM | 2 |
| 2023 | Optimizing Link-Identified Forwarding Framework in LEO Satellite NetworksabstractLow earth orbit (LEO) satellite networks have the potential to provide low-latency communication with global coverage. To unleash this potential, it is crucial to achieve efficient data delivery. In this paper, we analyze the topology characteristics of LEO satellite networks, and propose a source-route-style forwarding framework. Specifically, we leverage the deterministic neighbor relationship and identify all the unidi-rectional inter-satellite links (ISLs). Moreover, our framework utilizes the in-packet bloom filter (BF) to store the source-route-style forwarding information. This way, the source satellite could encode multiple ISL identifiers into the BF, which actually specifies the forwarding path. The intermediate satellites only need to check whether the outgoing ISLs are encoded and forward packets accordingly. Due to false positives caused by BF, the more ISLs are encoded at a time, the more redundant forwardings emerge. To reduce forwarding overhead, we take into account segment encoding, allowing the source and intermediate satellites to encode part of ISLs towards the destination. Overall, segment encoding seeks the right balance between forwarding overhead and encoding delay. We characterize a wide range of segment encoding policy in a unified framework, and derive the expected forwarding overhead in a closed-form. The segment encoding design is formulated as a binary non-linear programming, which is NP-hard. To overcome the challenge, we leverage its decomposable structure, and propose an efficient algorithm to solve it optimally. Finally, we validate our analytical results via packet-level experiments. Results also show that our proposed segment encoding policy significantly reduces the queuing delay compared to source encoding. Hefan Zhang 0002, Zhiyuan Wang 0004, Shan Zhang 0001, Qingkai Meng 0001, Hongbin Luo |
WiOpt | 4 |
| 2023 | Revisiting Congestion Detection in Lossless NetworksabstractCongestion detection is the cornerstone of end-to-end congestion control. Through in-depth observations and understandings, we reveal that existing congestion detection mechanisms in mainstream lossless networks (i.e., Converged Enhanced Ethernet and InfiniBand) are improper, due to failing to cognize the interaction between hop-by-hop flow controls and congestion detection behaviors in switches. We define the ternary states of switch ports and present Ternary Congestion Detection (TCD) for mainstream lossless networks. TCD utilizes the ON-OFF sending pattern and the feature of queue length evolutions to detect the transitions among ternary states. We also enable TCD under the practical multiple queues scenario by TCD-MQ. Testbed and extensive simulations demonstrate that TCD can detect congestion ports accurately and identify flows contributing to congestion as well as flows only affected by hop-by-hop flow controls. Meanwhile, we shed light on how to incorporate TCD with rate control. Case studies show that existing congestion control algorithms can achieve$3.3\times $and$2.0\times $better median and 99th-percentile FCT slowdown by combining with TCD. Qingkai Meng 0001, Fengyuan Ren |
IEEE/ACM Trans. Netw. | 2 |
| 2022 | CrossDBT: An LLVM-Based User-Level Dynamic Binary Translation Emulator
Wei Li 0262, Xiaohui Luo, Qingkai Meng 0001, Fengyuan Ren |
Euro-Par | 4 |
| 2022 | Demystifying and Mitigating TCP CappingabstractToday’s Internet user experience greatly depends on some user-perceived network metrics, such as throughput and latency. To improve these metrics, many Internet content providers build the content delivery network (CDN) to provide their services. Generally, CDNs adopt TCP as their transport protocol. A recent line of work improves TCP by proposing novel congestion control algorithms. However, we measure TCP performance in the production CDN and identify an interesting phenomenon termed TCP capping. When the flows experience TCP capping, the fixed-size receive window (rwnd) restricts these flows from fully utilizing network bandwidth. Through in-depth analysis, we demystify that the root cause of TCP capping is an inappropriate constraint on rwnd due to not considering the receiver’s processing capability. To mitigate it, this paper proposes a server-side scheme Apollo and a client-side scheme Artemis for Internet content providers and users, respectively. Apollo probes the receiver’s processing capability and assists the sender in packet sending. And Artemis adjusts the receive buffer in light of the receiver’s processing capability. In our evaluation, compared to vanilla TCP, TCP (w/ Apollo) and TCP (w/ Artemis) shorten flow completion time by up to 91.8% and 94.9%, respectively. Qingkai Meng 0001, Fengyuan Ren, Tong Zhang 0018, Danfeng Shan, Yajun Yang |
IWQoS | 1 |
| 2021 | RBA: Adaptive TCP Receive Buffer SizingabstractWith the rapid growth of hardware devices, a single host may have simultaneous connections that vary in network bandwidth and CPU processing capability as several orders of magnitude. State-of-art flow control mechanism, i.e., TCP auto-tuning, still needs to configure the maximum receive buffer, which cannot be applied to all connections in one host. In this paper, we reveal that improper receive buffer restrained by this configuration either (i) underutilizes the available network and CPU resources or (ii) occupies too much memory and then causes overall throughput collapse. To fully utilize resources with less memory occupancy, we present Receive Buffer Adaptive-regulating (RBA) algorithm, which regulates receive buffer according to the estimation of network bandwidth and receiver's processing capability. Testbed experiments show that RBA adapts to different scenarios and brings substantial performance improvement compared to TCP auto-tuning. Qingkai Meng 0001, Kun Qian 0017, Wenxue Cheng, Fengyuan Ren |
ISCC | 1 |
| 2021 | Lightning: A Practical Building Block for RDMA Transport ControlabstractRoCEv2 (RDMA over Converged Ethernet version 2) is the canonical method for deploying RDMA in Ethernet-based datacenters. Traditionally, RoCEv2 runs over the lossless network which is in turn achieved by enabling Priority Flow Control (PFC) within the network. However, with the scale of data center increases, PFC’s side effects, such as head-of-line blocking, congestion spreading, and PFC storms, are amplified. Datacenter operators can no longer tolerate these problems. They are seeking PFC alternatives for RDMA networks. Rather than aim at the lossless RDMA network, we instead handle packet loss effectively to support RDMA over Ethernet.In this paper, we propose Lightning, a switch building block to enhance RoCE’s simple loss recovery. Lightning enhances the switches to send loss notifications directly to the sources with high priority, thus informing sources as quickly as possible. Then, sources can retransmit packets sooner. By addressing challenges such as that shared buffer status is not available at ingress in modern switches, Lightning generates loss notification only when the expected packet is dropped and filters other unexpected packets at ingress, so as to avoid timeouts and prevent unnecessary congestion from unexpected packets. We implement Lightning on commodity programmable switches. In our evaluation, Lightning achieves up to 16.08× reduction of 99.9th percentile flow completion time compared to PFC, IRN and other alternatives. Qingkai Meng 0001, Fengyuan Ren |
IWQoS | 1 |
| 2021 | Congestion detection in lossless networksabstractCongestion detection is the cornerstone of end-to-end congestion control. Through in-depth observations and understandings, we reveal that existing congestion detection mechanisms in mainstream lossless networks (i.e., Converged Enhanced Ethernet and InfiniBand) are improper, due to failing to cognize the interaction between hop-by-hop flow controls and congestion detection behaviors in switches. We define ternary states of switch ports and present Ternary Congestion Detection (TCD) for mainstream lossless networks. Testbed and extensive simulations demonstrate that TCD can detect congestion ports accurately and identify flows contributing to congestion as well as flows only affected by hop-by-hop flow controls. Meanwhile, we shed light on how to incorporate TCD with rate control. Case studies show that existing congestion control algorithms can achieve 3.3x and 2.0x better median and 99th-percentile FCT slowdown by combining with TCD. Qingkai Meng 0001, Fengyuan Ren |
SIGCOMM | 3 |
| 2019 | Active and Adaptive Application-Level Flow Control for Latency Sensitive RPC ApplicationsabstractThe Remote Procedure Call (RPC) frameworks are widely deployed in industry. Applications supported by RPC frameworks are often latency-sensitive which strictly require to be responded before the deadline. For meeting this requirement, RPC frameworks adopt the application-level flow control mechanism. This mechanism gives an appropriate threshold determining the number of RPC requests that the server can process, thus avoids missing the deadline. However, this threshold at the application-level is a fixed empirical value so that it is hard to obtain respectable performance because an endpoint's processing capacity can take a huge quantity of values by varying workload and different hardware configurations. While other methods based on specialized transport protocols are adaptive, they will introduce extra costs for message reporting from server to client. Furthermore, adopting specialized transport protocols will also introduce extra transplanting efforts for TCP-based applications. In this paper, we provide an active and adaptive application level flow control mechanism at the client side. We first design an algorithm to find the appropriate threshold to achieve the desired response time. Then based on this algorithm, we control the threshold to bound the response time as expected. We implement our flow control mechanism using a memcached testbed. Experiments prove that our mechanism can accurately reduce the mean and 99th percentile response time by at least 71.3% and 69.4% respectively, while keeping a relatively high QPS. Furthermore, compared to static-threshold mechanism, our flow control mechanism is more efficient under low latency constraints. Jing Xie 0005, Wenxue Cheng, Tong Zhang 0018, Qingkai Meng 0001, Fengyuan Ren |
ICPADS | 4 |
| 2019 | Comparing Busy Poll Socket and NAPIabstractLow response time is the requirement for high performance computing applications. The network latency is an important influence factor on response time. Busy poll socket (BPS) is a new mechanism that can reduce network latency by introducing extra energy cost to busily poll RX queue. Unlike hardware dependent solutions such as RDMA, busy poll socket can improve performance without specialized hardware and transplanting applications. BPS now becomes mature and is supported by many off-the-shelf network interface cards. In this paper, we compare the mechanisms of BPS and NAPI. We make a comprehensive analysis about both BPS and NAPI mechanisms. We analyze which factors would influence latencies in BPS and NAPI modes respectively. We find out that BPS does not definitely outperform NAPI, and the reasons are explored in detail. Our findings can provide suggestions on the future deployment of BPS based applications. Jing Xie 0005, Qingkai Meng 0001, Xunli Fan, Niu Bo, Fengyuan Ren |
ICPADS | 3 |
| 2017 | SoftRDMA: Rekindling High Performance Software RDMA over Commodity EthernetabstractRecent academic and industrial work is exploring the challenges of using RDMA over Ethernet, to support highly reliable, latency-sensitive services in today's datacenters. Previous work on the high-speed packet I/O like netmap, DPDK, etc., and high-performance user-level stacks like mTCP, IX etc., rekindles our inspirations to implement a high-performance software RDMA over commodity Ethernet devices. Mao Miao, Fengyuan Ren, Xiaohui Luo, Jing Xie 0005, Qingkai Meng 0001, Wenxue Cheng |
APNet | 5 |