VLDB 2026 Research / reviewers in the wild / expert
Kunling He
dblp:278/2703
· DBLP profile ↗
9ranked-venue papers
1as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 6 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ECOTE: Priority-Aware Optical Restoration for WAN Traffic EngineeringabstractFiber cuts are among the most common and disruptive failures in cloud networks. They can prevent cloud providers from maintaining committed service availability, causing Service Level Agreement (SLA) violations that directly translate into monetary penalties. Existing traffic engineering (TE) approaches enhance failure resilience, and recent systems further incorporate optical restoration to recover lost bandwidth after failures. However, they still treat services largely uniformly and optimize primarily for network throughput rather than the economic impact of heterogeneous SLA penalties. In this paper, we present ECOTE, the first priority-aware TE system with optical restoration that explicitly minimizes revenue loss. Specifically, ECOTE introduces a new optical restoration formulation with a dedicated capacity restoration solver to compute an optimal restoration plan that maximizes restorable capacity under physical constraints. ECOTE also designs a priority-aware TE algorithm that allocates the restored bandwidth capacity according to SLA penalties, thereby reducing monetary cost. We evaluate ECOTE using a production-level WAN testbed and through large-scale simulations. The testbed evaluation demonstrates ECOTE achieves zero loss for high-priority services and more than 10× revenue loss reduction compared to state-of-the-art. Our large-scale simulation results show that ECOTE can support at least 2.5× and 2.0× more demand for different high priority services compared to the state-of-the-art solutions. Meanwhile, ECOTE reduces the revenue loss by at least an order of magnitude less than existing solutions. Kunling He, Ran Shu 0001, Jilong Wang 0001, Congcong Miao |
EuroSys | 2 |
| 2026 | Balancing and Beyond: Communication-Centric Optimizations in Expert ParallelismabstractThe Mixture-of-Experts (MoE) architecture scales large language models (LLMs) to trillions of parameters by activating only a small subset of experts per token. In practice, MoE inference is commonly deployed with Expert Parallelism (EP), which places whole experts on different GPUs to preserve kernel efficiency. However, production EP deployments often suffer from two bottlenecks: (1) expert workload imbalance, which creates computation and communication stragglers, and (2) communication inefficiency, where inter-GPU transfers dominate latency even after balancing. We present EPIC, an experience-driven EP inference system that addresses these issues progressively for real deployments. EPIC mitigates imbalance via performance-aware expert migration and runtime expert activation, and then improves communication with topology-adaptive transport kernels and fine-grained computation-communication overlap. EPIC has been deployed at scale across O(10K) GPUs in our online inference service for both open-source models (e.g., Qwen3-Coder and DeepSeek-R1) and internal models, reducing communication time and per-token latency by up to 40% and 21%, respectively. Jiamin Cao, Qingxu Li, Yaozhong Liu, Shangfeng Shi, Kunling He, Ennan Zhai, Jianbo Dong, Binzhang Fu, Dennis Cai |
SIGCOMM | 10 |
| 2025 | Flexnetic: Cost-Effective and Smooth Evolution of Optical BackboneabstractThe increasing traffic on WANs due to growing number of applications imposes significant strain on the infrastructure of cloud service providers. The primary expense in augmenting network capacity entails high costs associated with procuring expensive transponders that facilitate inter-regional optical signal transmission. The advent of spacing variable transponders facilitates cost-effective network upgrades and increased transmission efficiency, however, implementing such upgrades in one step may incur high costs and service disruptions. Achieving a cost-effective and smooth network upgrade presents considerable challenges in identifying critical IP links and minimizing costs within existing architecture. We introduce Flexnetic, a planning tool which utilizes a hybrid approach of both modern and legacy transponders, along with establishment of optical bypass, to accommodate the escalating traffic demands while minimizing the costs during network upgrades. Flexnetic incorporates two novel algorithms: a NLP model to maximize IP link capacity utilization, and an MIP model for efficient IP layer implementation at optical layer, emphasizing reuse of existing transponder and minimizes new device requirement. Our simulation of upgrade plans on common WAN topologies revealed superior performance, with up to 91.9% cost savings and 2.33× capacity increase over existing state-of-the-art solutions, highlighting Flexnetic’s potential for cost-efficient and capacity-optimized network upgrades. Congcong Miao, Kunling He, Jilong Wang 0001 |
ICCCN | 4 |
| 2025 | SimAI: Unifying Architecture Design and Performance Tuning for Large-Scale Large Language Model Training with Scalability and Precision
Xizheng Wang, Qingxu Li, Yichi Xu, Dan Li 0001, Li Chen 0008, Heyang Zhou, Linkang Zheng, Yikai Zhu, Yang Liu 0245, Kun Qian 0021, Kunling He, Ennan Zhai, Dennis Cai, Binzhang Fu |
NSDI | 14 |
| 2025 | Alibaba Stellar: A New Generation RDMA Network for Cloud AIabstractThe rapid adoption of Large Language Models (LLMs) in cloud environments has intensified the demand for high-performance AI training and inference, where Remote Direct Memory Access (RDMA) plays a critical role. However, existing RDMA virtualization solutions, such as Single-Root Input/Output Virtualization (SR-IOV), face significant limitations in scalability, performance, and stability. These issues include lengthy container initialization times, hardware resource constraints, and inefficient traffic steering. To address these challenges, we propose Stellar, a new generation RDMA network for cloud AI. Stellar introduces three key innovations: Para-Virtualized Direct Memory Access (PVDMA) for on-demand memory pinning, extended Memory Translation Table (eMTT) for optimized GPU Direct RDMA (GDR) performance, and RDMA Packet Spray for efficient multi-path utilization. Deployed in our large-scale AI clusters, Stellar spins up virtual devices in seconds, reduces container initialization time by 15 times, and improves LLM training speed by up to 14%. Our evaluations demonstrate that Stellar significantly outperforms existing solutions, offering a scalable, stable, and high-performance RDMA network for cloud AI. Menglei Zheng, Binbin Liao, Suwei Xu, Yongjia Mo, Qinghua Peng, Jilie Luo, Qingxu Li, Zishu Wang, Jianbo Dong, Kunling He, Sheng Cheng 0002, Jiamin Cao, Hairong Jiao, Lingjun Zhu, Yiquan Chen, Wei Wang 0030, Shuhong Zhu, Xingru Li, Qiang Wang 0022, Wei Lin 0016, Ennan Zhai, Jiesheng Wu, Qiang Liu 0036, Binzhang Fu, Dennis Cai |
SIGCOMM | 18 |
| 2023 | GapReplay: A High-Accuracy Packet ReplayerabstractNetwork traffic has become increasingly complicated and diverse with the continuing development of the Internet. It is challenging for the traffic generators to construct a large variety of traffic while maintaining the similarity to the real traffic. Given that real traffic from the production network is handy to retrieve, replaying traces from the production network can maximize the authenticity of traffic, contributing to experiments to evaluate the effectiveness of algorithms. Recently, many pieces of research in the networking field, especially algorithms based on machine learning (ML), are extremely sensitive to statistical characteristics of the traffic, such as ML-based traffic classification, traffic detection, traffic prediction, etc. However, there are only a few available statistics for the growing encrypted traffic, such as packet size and inter-arrival time (IAT). To replay traffic accurately, in this paper, we propose the GapReplay, a packet replayer that can remain identical with the original nanosecond-precision pcap trace in packet contents and achieve high accuracy in timestamps. GapReplay reaches line rate while transmitting packets by appending and extending packets, which will be dropped and truncated by a programmable device, to guarantee satisfying accuracy. We implement our mechanism on top of DPDK and P4. The evaluation results demonstrate that GapReplay can achieve a nanosecond-level accuracy, much better than the state-of-the-art such as MoonGen and tcpreplay, where the best of them is at least 1000 times less accurate than our mechanism. Shuanghong Yu, Han Zhang 0009, Kunling He, Xingkun Yao |
ICC | 4 |
| 2023 | FlexWAN: Software Hardware Co-design for Cost-Effective and Resilient Optical BackbonesabstractThe rising demand for WAN capacity driven by the rapid growth of inter-data center traffic poses new challenges for costly optical networks. Today cloud providers rely on fixed optical backbones, where all hardware devices operate on a rigid spectrum grid, leading to the waste of expensive optical resources and subpar performance in handling failures. In this paper, we introduce FlexWAN, a novel flexible WAN infrastructure designed to provision cost-effective WAN capacity while ensuring resilience to optical failures. FlexWAN achieves this by incorporating spacing-variable hardware at the optical layer, enabling the generated wavelength to optimize the utilization of limited spectrum resources for the WAN capacity. The configuration of spacing-variable hardware in a multi-vendor optical backbone presents challenges related to spectrum management. To address this, FlexWAN leverages a centralized controller to achieve coordinated control of network-wide optical devices in a vendor-agnostic manner. Moreover, the flexibility at the optical layer introduces new algorithmic problems. FlexWAN formulates the problem of provisioning WAN capacity with the goal of minimizing hardware costs. We evaluate the system performance in production and share insights from years of production experience. Compared to existing optical backbones, FlexWAN can save at least 57% of transponders and reduce 36% of spectrum usage while continuing to meet up to 8× the present-day demands using existing hardware and fiber deployments. FlexWAN further incorporates failure resilience that revives 15% more bandwidth capacity in the overloaded optical backbone. Congcong Miao, Zhizhen Zhong, Ying Zhang 0022, Kunling He, Fangchao Li, Minggang Chen, Xiang Li 0223, Zekun He, Xianneng Zou, Jilong Wang 0001 |
SIGCOMM | 4 |
| 2022 | Multi-hop Precision Time Protocol: an Internet Applicable Time Synchronization SchemeabstractPrecise time synchronization is essential for 5G systems, data centers, industrial automation systems, military fields, and more. Although the precision of IEEE 1588 PTP can achieve sub-hundred-nanosecond accuracy, it works only when being deployed hop-by-hop within a LAN with limited range. Hop-by-hop deployment leads to high deployment costs and makes it inapplicable over the Internet. In this paper, we propose the multi-hop precision time protocol (M-PTP), a high-precision and low-cost time synchronization protocol, which does not require hop-by-hop deployment, and no special functions need to be added to the network relay devices such as routers and switches. M-PTP leverages two key ideas. First, to mitigate the "asymmetry in forward delay and reverse delay" problem, SVM-based delay estimation is used to calculate the distribution of positive and negative random delays, then L-estimator is leveraged to estimate the time offset. Second, based on time offset, M-PTP exploits loop effect optimization among nodes. We implemented the protocol and tested its performance on variance hops under different traffic conditions and CPU loads. The experimental results show that M-PTP can achieve a precision of 11.61ns at 5 hops, which is approximately 3 times the precision of HUYGENS and approximately 30 times the precision of PTP. Kunling He, Changqing An, Hui Wang 0011, Tianshu Li, Linmei Zu |
NOMS | 1 |
| 2020 | A transfer learning method using speech data as the source domain for micro-Doppler classification tasks
Kunling He, Danlei Xu, Dingli Luo |
Knowl. Based Syst. | 2 |