Ke Li 0020

dblp:75/6627-20 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0002-2189-7967ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 12 · 2 first-author · 11 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Approximate Gradient Synchronization With Adaptive Quantized Gradient Broadcast
Shouxi Luo, Ke Li 0020, Huanlai Xing
Future Gener. Comput. Syst.3
2026 Maximizing the benefits of in-network aggregation with joint job placement and routing control
Shouxi Luo, Huanlai Xing, Ke Li 0020, Bo Peng 0006
Future Gener. Comput. Syst.4
2025 Efficient In-Network Aggregation With Adaptive Quantization
Zhongxu Su, Shouxi Luo, Ke Li 0020, Huanlai Xing, Bo Peng 0006
APNet3
2025 Dynamic Online Resource Allocation for Synchronization, Retraining, and Inference in Digital Twin Network
abstract
With the advancement of Intelligent Transportation Systems (ITS), Digital Twin (DT) technology has been widely applied to tasks such as traffic flow modeling and autonomous driving assistance. However, traditional standalone DT models often suffer from weak generalization ability and potential risks of privacy leakage in data-drifting scenarios. To address these challenges, a novel FL-DTN architecture for vehicular networks is proposed by integrating Federated Learning (FL) with Digital Twin Networking (DTN), aiming to preserve data privacy and enhance generalization ability of digital twin models. Specifically, an online resource allocation algorithm, Online Resource Allocation for Synchronization, Retraining and Inference (ORASRI), is designed to dynamically balance the resource allocation among digital twin synchronization, retraining, and inference for Vehicle Digital Twins (VDTs), while adapting to data drift under constrained resource conditions. In addition, a Teacher-Student collaborative mechanism is introduced to improve inference accuracy while reducing resource consumption. Experiments on the MNIST-C dataset show that ORASRI improve 4.6% and 10.3% Inference accuracy in non-FL and FL data-drifting scenarios, respectively.
Ke Li 0020, Weichen Tian, Penglin Dai, Shouxi Luo, Huanlai Xing
GLOBECOM1
2025 AoI-Error-Aware Data Synchronization for Vehicular Digital Twin
abstract
Synchronization of vehicular digital twin (VDT) state data is essential for maintaining the accuracy of digital twin models. Existing studies show that VDT data synchronization typically requires a substantial amount of bandwidth and frequent data exchanges. However, real-world vehicle state data are prone to noise interference and bandwidth constraints, causing synchronization errors and significantly degrading VDT accuracy. To evaluate the impact of noise interference on VDT accuracy, we model vehicle state evolution as a discrete-time wiener process and employ a Kalman filter for optimal state estimation. By further analyzing the relationship between estimation error and update timeliness, we find that weighted scheduling based on Age of Information (AoI) effectively suppresses error accumulation and improves synchronization performance. Then, we propose an aoi error-aware scheduling mechanism maximum weighted noise age (MWNA), within a cloud-edge collaborative VDT framework, MWNA dynamically evaluates each vehicle’s state update and prioritizes the transmissions that yield the greatest reduction in estimation error. Compared with baselines, the MWNA policy achieves up to 25–40% lower average error across a range of correlation settings, particularly under high noise correlation and large-scale vehicular scenarios.
Ke Li 0020, Xinbang Zhang, Haojun Huang, Shouxi Luo, Huanlai Xing
GLOBECOM1
2025 Capacity-Achieving Coding Schemes of Gaussian Finite-State Markov Wiretap Channels With Delayed Feedback
Dengfeng Xia, Ke Li 0020, Peng Xu 0002, Bin Dai 0003, Liuguo Yin
IEEE Trans. Inf. Forensics Secur.2
2025 Joint Optimization of Device Placement and Model Partitioning for Cooperative DNN Inference in Heterogeneous Edge Computing
abstract
EdgeAI represents a compelling approach for deploying DNN models at network edge through model partitioning. However, most existing partitioning strategies have primarily concentrated on homogeneous environments, neglecting the effect of device placement and their inapplicability to heterogeneous settings. Moreover, these strategies often rely on either data parallelism or model parallelism, each presenting its own limitations, including data synchronization and communication overhead. This paper aims at enhancing inference performance through a pipeline system of devices through leveraging both parallel and sequential relationships among them. Accordingly, the problem of Multi-Device Cooperative DNN Inference is formulated by optimizing both device placement and model partitioning, taking into account the unique characteristics of heterogeneous edge resources and DNN models, with the goal of maximizing throughput. To this end, we propose an evolutionary device placement technique to determine the pipeline stage of devices by enhancing a variant of particle swarm optimization. Subsequently, an adaptive model partitioning strategy is developed by combining intra-layer and inter-layer model partitioning based on dynamic programming and the input-output mapping of DNN layers, respectively, to accommodate edge resource limitations. Finally, we construct a simulation model and a prototype, and the extensive results demonstrate that our proposed algorithm outperforms current state-of-the-art algorithms.
Penglin Dai, Biao Han 0001, Ke Li 0020, Xincao Xu, Huanlai Xing, Kai Liu 0001
IEEE Trans. Mob. Comput.3
2025 Efficient Parameter Synchronization for Peer-to-Peer Distributed Learning With Selective Multicast
abstract
Recent advances in distributed machine learning show theoretically and empirically that, for many models, provided that workers will eventually participate in the synchronizations,$i)$the training still converges, even if only$p$workers take part in each round of synchronization, and$ii)$a larger$p$generally leads to a faster rate of convergence. These findings shed light on eliminating the bottleneck effects of parameter synchronization in large-scale data-parallel distributed training and have motivated several optimization designs. In this paper, we focus on optimizing the parameter synchronization forpeer-to-peerdistributed learning, where workers broadcast or multicast their updated parameters to others for synchronization, and proposeSelMcast, a suite of expressive and efficient multicast receiver selection algorithms, to achieve the goal. Compared with the state-of-the-art (SOTA) design, which randomly selects exactly$p$receivers for each worker’s multicast in a bandwidth-agnostic way,SelMcastchooses receivers based on the global view of their available bandwidth and loads, yielding two advantages, i.e., accelerated parameter synchronization for higher utilization of computing resources and enlarged average$p$values for faster convergence. Comprehensive evaluations show thatSelMcastis efficient for both peer-to-peer Bulk Synchronous Parallel (BSP) and Stale Synchronous Parallel (SSP) distributed training, outperforming the SOTA solution significantly.
Shouxi Luo, Pingzhi Fan, Ke Li 0020, Huanlai Xing, Long Luo, Hong-Fang Yu
IEEE Trans. Serv. Comput.3
2024 ADCC: AoI-aware Decentralized Congestion Control in Cooperative Perception System
abstract
Cooperative perception, based on vehicle-to-everything (V2X) communication technology, is a promising solution for connected and automated vehicles (CAVs) to improve their perception capabilities in intelligent transportation systems. The frequency of message transmission in cooperative perception among mobile vehicles plays a crucial role, as it directly impacts communication efficiency, perception accuracy, system response speed, and the safety of real-time applications. Higher transmission frequencies can provide more timely and rich sensory information. This also implies higher consumption of communication resources. However, in a dynamic and complex environment, it is difficult to quantitatively control the message transmission frequency so as to improve the utilization of limited communication resources, ensure the timeliness of perception messages, and maintain fast convergence. To address challenges of timeliness of messages, a timeliness performance metric age of information (AoI) is introduced to control message transmission frequency. This paper deduced the average AoI of the system and designed an AoI-aware decentralized congestion control (ADCC) algorithm for the V2X-based cooperative perception system. Simulation results show that the ADCC algorithm outperforms the classic congestion control algorithm linear adaptive message rate (LIMERIC) in terms of channel utilization and throughput. Specifically, AoI decreased by 47.3%, channel utilization increased by 6.5% and throughput increased by 48.6%.
Ke Li 0020, Haojun Huang, Shouxi Luo, Huanlai Xing
HPCC1
2024 Towards Optimal Topology-Aware AllReduce Synthesis
abstract
In this work, we propose TARS, a Topology-aware AllReduce algorithm Synthesizer, to generate optimal execution plans for AllReduce workloads over arbitrary interconnection network structures. Distinguished from existing topology-aware synthesizers that formulate the two stages of AllReduce (e.g., ReduceScatter-then-AllGather, or Reduce-then-Broadcast) separately, the power of TARS stems from employing a comprehensive Integer Quadratic Programming (IQP) model to formulate the entire workflow precisely. Preliminary studies confirm that, compared with the state-of-the-art scheme, TARS could significantly reduce the completion time of AllReduce.
Wenhao Lv, Shouxi Luo, Ke Li 0020, Huanlai Xing
IWQoS3
2024 Releasing the Power of In-Network Aggregation With Aggregator-Aware Routing Optimization
abstract
By offloading partial of the aggregation computation from the logical central parameter servers to network devices like programmable switches, In-Network Aggregation (INA) is a general, effective, and widely used approach to reduce network load thus alleviating the communication bottlenecks suffered by large-scale distributed training. Given the fact that INA would take effects if and only if associated traffic goes through the same in-network aggregator, the key to taking advantage of INA lies in routing control. However, existing proposals fall short in doing so and thus are far from optimal, since they select routes for INA-supported traffic without comprehensively considering the characteristics, limitations, and requirements of the network environment, aggregator hardware, and distributed training jobs. To fill the gap, in this paper, we systematically establish a mathematical model to formulate i) the up-down routing constraints of Clos datacenter networks, ii) the limitations raised by modern programmable switches’ pipeline hardware structure, and iii) the various aggregator-aware routing optimization goals required by distributed training tasks under different parallelism strategies. Based on the model, we develop ARO, an Aggregator-aware Routing Optimization solution for INA-accelerated distributed training applications. To be efficient, ARO involves a suite of search space pruning designs, by using the model’s characteristics, yielding tens of times improvement in the solving time with trivial performance loss. Extensive experiments show that ARO is able to find near-optimal results for large-scale routing optimization in tens of seconds, achieving$1.8\sim 4.0\times $higher throughput than the state-of-the-art solution.
Shouxi Luo, Ke Li 0020, Huanlai Xing
IEEE/ACM Trans. Netw.3
2024 Efficient Cross-Cloud Partial Reduce With CREW
abstract
By allowing$p$out of$n$workers to conductall reduceoperations among them for a round of synchronization,partial reduce, a promising partially-asynchronous variant ofall reduce, has shown its power in alleviating the impacts of stragglers for iterative distributed machine learning (DML). Currentpartial reducesolutions are mainly designed for intra-cluster DML, in which workers are networked with high-bandwidth LAN links. Yet no prior work has looked into the problem of how to achieve efficientpartial reducefor cross-cloud DML, where inter-worker connections are with scarcely-available capacities. To fill the gap, in this paper, we proposeCREW, a flexible and efficient implementation ofpartial reducefor cross-cloud DML. At the high level,CREWis built upon the novel design of employing all active workers along with their internal connection capacities to execute the involved communication and computation tasks; and at the low level,CREWemploys a suite of algorithms to distribute the tasks among workers in a load-balanced way, and deal with possible outages of workers/connections, and bandwidth contention. Detailed performance studies confirm that,CREWnot only shortens the execution of eachpartial reduceoperation, outperforming existing communication schemes such as PS, Ring,TopoAdopt, and BLINK greatly, but also significantly accelerates the training of large models, up to$15\times$and$9\times$, respectively, when compared with the all-to-all direct communication scheme andoriginal partial reducedesign.
Shouxi Luo, Renyi Wang, Ke Li 0020, Huanlai Xing
IEEE Trans. Parallel Distributed Syst.3
2024 Joint Optimization for Quality Selection and Resource Allocation of Live Video Streaming in Internet of Vehicles
abstract
Live Video Streaming (LVS) services are critical in supporting real-time applications in Internet of Vehicles (IoV) by transmitting real-time generated video content from streaming server to vehicles. Due to restricted spectrum resources and high vehicle mobility, LVS suffers from notable performance degradation. Moreover, existing strategies such as buffer size control and edge caching, are designed for video-on-demand service, which is ineffective for LVS in IoV. Accordingly, we investigate the problem of LVS-IoV by synthesizing multicasting and Scalable Video Coding-based encoding with the goal of maximizing Quality of Experience (QoE), which is defined as the weighted sum of video quality, rebuffering time, and quality variation. The LVS-IoV is decoupled into three sub-problems: vehicle grouping, quality selection, and resource allocation. Firstly, we propose a K-means-based vehicle grouping method that considers geographical distribution, velocity, and dynamic channels. Secondly, we determine the quality selection of each group based on the Value Decomposition Network for maximizing overall video quality. This network utilizes global value function decomposition and centralized training to achieve fast convergence, followed by distributed execution. Lastly, we propose a sub-gradient algorithm to achieve optimal resource allocation. We build simulation model and perform extensive evaluation, which demonstrates its superiority compared to other competitive methods.
Penglin Dai, Meiting Wu, Ke Li 0020, Xiao Wu 0001, Yan Ding 0002
IEEE Trans. Serv. Comput.3
2022 Approximate Gradient Synchronization with AQGB
abstract
No abstract available.
Shouxi Luo, Ke Li 0020, Huanlai Xing
APNet3
2022 AoI-aware Data Propagation in Edge Computing Assisted NR-V2X Networks
abstract
Intelligent Transportation System (ITS) applications in vehicle networks often use latency, reliability, and throughput as performance metrics. However, these traditional metrics fail to indicate the freshness of data. ”Age-of-information” (AoI) has been emphasized as a critical requirement for real-time applications as a new performance metric. Traditional broadcast-based data propagation tends to cause congestion because it needs a lot of vehicles to participate in the data transmission, resulting in large AoI. With the help of mobile edge computing, this paper proposes a new unicast-based data propagation approach to reduce AoI. In this scenario, an edge computing server computes and determines vehicle nodes for the purpose of propagating time-sensitive data via V2V communication. The goal of this research is to provide an AoI-aware unicast data propagation method that optimizes the initial vehicle nodes selection and design scheduling rules for maintaining information fresh under a coverage rate restriction (corresponding to emergency level). Simulation results show that the proposed unicast-based method has lower AAoI and PAoI than that of the traditional broadcast-based method.
Kexun Chen, Ke Li 0020
APNet3
2022 Efficient Partial Reduce Across Clouds
abstract
No abstract available.
Renyi Wang, Shouxi Luo, Ke Li 0020, Huanlai Xing
APNet3
2022 Fast Parameter Synchronization for Distributed Learning with Selective Multicast
abstract
Recent advances in distributed machine learning show theoretically and empirically that, for many models, provided workers would participate in the synchronizations eventually, i) the training still converges, even if only p workers take part in each round of synchronization, and ii) a larger p generally leads to a faster rate of convergence. These findings shed light on eliminating the bottleneck effects of parameter synchronization in large-scale data-parallel distributed training, having motivated several optimization designs.In this paper, we focus on optimizing the parameter synchronization for peer-to-peer distributed learning, in which workers generally broadcast or multicast their updated parameters to others for synchronization, and propose SELMCAST, an expressive and Pareto-optimal multicast receiver selection algorithm, to achieve the goal. Compared with the state-of-the-art design that randomly selects exactly p receivers for each worker’s multicast in a bandwidth-agnostic way, SELMCAST chooses receivers based on the global view of their available bandwidth and loads, yielding two advantages. Firstly, it could optimize the bottleneck sending rate, thus cutting down the time cost of parameter synchronization. Secondly, when more than p receivers are with sufficient bandwidth, they would be selected as many as possible, bringing benefits to the convergence of training. Extensive evaluations show that SELMCAST is efficient and always achieves near-optimal performance.
Shouxi Luo, Pingzhi Fan, Ke Li 0020, Huanlai Xing, Long Luo, Hong-Fang Yu
ICC3
2022 Poster: Selective Reduce for Heterogeneous Distributed Training
abstract
To improve the performance of partial reduce on synchronizing models for heterogeneous data-parallel distributed training, we explore the idea of selective waiting and worker selection to propose the flexible solution of selective reduce. Our preliminary study shows that with progress- and bandwidth-aware decisions, the proposed partial reduce outperforms the original partial reduce significantly, in terms of both the average synchronization scales and completion times.
Shouxi Luo, Ke Li 0020, Huanlai Xing
ICNP3
2022 Offloading dependent tasks in multi-access edge computing: A multi-objective reinforcement learning approach
Fuhong Song, Huanlai Xing, Xinhan Wang, Shouxi Luo, Penglin Dai, Ke Li 0020
Future Gener. Comput. Syst.6
2022 SelfMatch: Robust semisupervised time-series classification with self-distillation
abstract
Over the years, a number of semisupervised deep-learning algorithms have been proposed for time-series classification (TSC). In semisupervised deep learning, from the point of view of representation hierarchy, semantic information extracted from lower levels is the basis of that extracted from higher levels. The authors wonder if high-level semantic information extracted is also helpful for capturing low-level semantic information. This paper studies this problem and proposes a robust semisupervised model with self-distillation (SD) that simplifies existing semisupervised learning (SSL) techniques for TSC, called SelfMatch. SelfMatch hybridizes supervised learning, unsupervised learning, and SD. In unsupervised learning, SelfMatch applies pseudolabeling to feature extraction on labeled data. A weakly augmented sequence is used as a target to guide the prediction of a Timecut-augmented version of the same sequence. SD promotes the knowledge flow from higher to lower levels, guiding the extraction of low-level semantic information. This paper designs a feature extractor for TSC, called ResNet–LSTMaN, responsible for feature and relation extraction. The experimental results show that SelfMatch achieves excellent SSL performance on 35 widely adopted UCR2018 data sets, compared with a number of state-of-the-art semisupervised and supervised algorithms.
Huanlai Xing, Zhiwen Xiao, Dawei Zhan, Shouxi Luo, Penglin Dai, Ke Li 0020
Int. J. Intell. Syst.6
2020 A Popularity- and Mobility-Aware Multi-layer Caching with Feedback Mechanism for Highway Vehicular Networks
abstract
With the explosion of the video content industry, the demand for large file downloads is increasing rapidly. Different from text-based contents, video contents require lossless transmission with an ordered constraint to ensure the best utility. Meanwhile, providing low latency and stable downloads under high mobility scenarios, such as vehicle users in a highway, is one of the major challenges in wireless communications. In this work, a caching strategy is adopted to address the problem of downloading large order-sensitive contents, supporting user mobility. Particularly, a multi-layer caching framework is proposed based on a simple yet efficient adaptive chunk-based caching (ACBC) algorithm. The multi-layer caching framework increases the flexibility in deploying file chunks and shortens the distance between user and content with affordable memory cost, while ACBC algorithm introduces a feedback mechanism to provide stable and high-efficiency chunk deployment scheme. The cooperation of the two ensures reliable downloads for high-mobility users with minimal access delay and caching space requirements. Simulation results indicate that our solution achieves a very small content access delay with a low cache space occupation compared with other mobility-based caching algorithms (such as MAP and RICH).
Ke Li 0020, Zhaoyan Lyu, Heng Liu 0009, Pingzhi Fan
VTC Fall1
2020 Efficient File Dissemination in Data Center Networks With Priority-Based Adaptive Multicast
abstract
In today's data center networks (DCN), cloud applications commonly disseminate files from a single source to a group of receivers for service deployment, data replication, software upgrade, etc. For these group communication tasks, recent advantages of software-defined networking (SDN) provide bandwidth-efficient ways-they enable DCN to establish and control a large number of explicit multicast trees on demand. Yet, the benefits of data center multicast are severely limited, since there does not exist a scheme that could prioritize multicast transfers respecting the performance metrics wanted by today's cloud applications, such as pursuing small mean completion times or meeting soft-time deadlines with high probability. To this end, we propose PAM (Priority-based Adaptive Multicast), a preemptive, decentralized, and ready-deployable rate control protocol for data center multicast. At the core, switches in PAM explicitly control the sending rates of concurrent multicast transfers based on their desired priorities and the available link bandwidth. With different policies of priority generation, PAM supports a range of scheduling goals. We not only prototype PAM upon the emerged P4-based programmable switch with novel approximation designs, but also evaluate its performance with ns3-based extensive simulations. Results imply that PAM is ready-deployable; it converges very fast, has negligible impacts on coexisting TCP traffic, and always performs near-optimal priority-based multicast scheduling.
Shouxi Luo, Hong-Fang Yu, Ke Li 0020, Huanlai Xing
IEEE J. Sel. Areas Commun.3