VLDB 2026 Research / reviewers in the wild / expert
Qing Li 0028
dblp:181/2689-28
· DBLP profile ↗
19ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0002-1772-9194ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 11 · 4 first-author · 11 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient and Effective Personalized In-Context Learning for On-Device Large Model ServicesabstractRecently, on-device large models (e.g., 7B parameter LLMs and LVLMs) are increasingly deployed in real-world services such as mobile assistants, edge computing, and IoT systems, where low latency and resource efficiency are critical. Since large models trained on general-domain data are typically suboptimal for downstream services, In-Context Learning (ICL) technique is widely used to enhance their task-specific service performance without extra tuning. However, effective ICL methods generally necessitate a sufficient amount of supervised data to provide abundant service-related information, which usually leads to a long context (i.e., more texts or images input) for models to process. This long context issue can significantly degrade the real-time inference efficiency and also impact the information utilization effectiveness, particularly for on-device large models with limited long context modeling capacities. To tackle this challenge, we propose a Personalized Knowledge Refinement (PKR) framework to achieve efficient and effective ICL for on-device large models. Specifically, we first introduce a personalized knowledge extraction module, which analyzes the inference behaviors of the target model on small-scale supervised data and then convert these behavioral patterns into personalized service-specific knowledge as the instruction context. Furthermore, we propose an adaptive knowledge filtration mechanism to model the informativeness of the extracted knowledge and eliminate the redundant ones, further improving the knowledge encoding efficiency per unit of context length. Experiments based on 14 benchmarks spanning both textual and visual task services, and 4 large models, demonstrated that PKR consistently improves task accuracy while reducing context length and inference latency, making it a practical solution for real-world on-device services. Codes are released athttps://github.com/wanghl21/PKR. Huili Wang 0001, Yuanhong Huang, Zhiyang Hu, Qing Li 0028, Pingyi Fan, Yongfeng Huang 0001, Shangguang Wang, Tao Qi 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2025 | Demo: Emulating Space Computing Networks with RHONEabstractThe rapid advancement in satellite technology with the adoption of commercial off-the-shelf (COTS) devices and satellite constellation networking has given rise to Space Computing Networks (SCNs). While SCN research is typically conducted on experimental platforms due to high operational costs, the unique challenges of SCNs require special consideration. In this demo, we introduce Rhone, an emulator that bridges these gaps by achieving both satellite- and constellation-level fidelity (the accurate replication of satellite and constellation states, including power, thermal, and network conditions, as well as application performance) while ensuring usability. Rhone adopts a two-phase approach: i) an offline phase builds power, thermal, orbit, network, and computation models using real satellite telemetries and hardware-in-the-loop chip mirroring, and ii) an online phase executes container-based emulation integrated with these models. Evaluation shows Rhone's power and computation model errors under 5% and thermal model errors within 1.3 – 2.5°C. Liying Wang 0011, Qing Li 0028, Shangguang Wang, Xuanzhe Liu, Chenren Xu |
MobiCom | 3 |
| 2025 | Emulating Space Computing Networks with RHONE
Liying Wang 0011, Qing Li 0028, Zhaofeng Luo, Shangguang Wang, Xuanzhe Liu, Chenren Xu |
USENIX ATC | 2 |
| 2025 | From Earth to Orbit: Launch Sequence Optimization for LEO Mega-ConstellationsabstractThe recent emergence of Low Earth Orbit (LEO) mega-constellations, designed for high-speed broadband connections with low latency, has introduced new deployment challenges. Efficient launch sequence planning is crucial for rapid service rollout, performance enhancement, and service promotion. However, existing research predominantly focuses on the design and performance analysis of fully-deployed constellations and overlooks the evolving process from a partially-deployed constellation to a fully-deployed one. This paper explores the launch sequence optimization problem for mega-constellations, tailored to expedite service delivery and adapt to changing performance demands. To this end, (1) we identify critical network performance metrics for the constellation evolving process and construct a simulation toolchain capable of simulating and evaluating these metrics for any potential partially-deployed constellation. (2) Drawing upon three key observations on network availability, the number of visible satellites, and latency, we propose an algorithm that can construct a launch sequence for an arbitrary mega-constellation topology. Evaluation results show that this algorithm enables the early provision of services and maximizes network performance gains at each launch batch while catering to different user demands. For instance, our algorithm can achieve network performance nearly equivalent to that of Starlink when it initiated its service, without losing redundancy, while using 55% fewer satellites. Qing Li 0028, Chenren Xu, Mengwei Xu 0001, Shangguang Wang, Gang Huang 0001, Xuanzhe Liu |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | FaaSPR: Latency-Oriented Placement and Routing Optimization for Serverless Workflow ProcessingabstractWorkflow processing enhances the applicability of serverless computing while retaining the characteristics of fine-grained resource management and elastic scalability. However, current serverless platforms lack targeted optimization of placement and routing strategies for workflow processing, leading to high overheads due to inter-server data transmission, instance cold starts, and function request queuing. We propose FaaSPR, a serverless scheduling system that exploits placement and routing optimizations to minimize workflow processing latency. FaaSPR groups instances with potential data transmission and proportionally distributes groups with heterogeneous instances across multiple servers, taking into account resource constraints and historical placement traces. This method addresses the issues of poor scalability and frequent instance migrations in existing solutions. Utilizing a routing algorithm based on multi-stage linear programming, FaaSPR minimizes cross-server data transmission within and between instance groups while ensuring load balancing among instances. Experiments show that, compared to the state-of-the-art solution FaaSFlow, FaaSPR decreases the average and 99th percentile tail latency by up to 68.03% and 93.46%, respectively. Additionally, reducing workflow processing latency leads to up to 46.18% decrease in resource consumption for FaaS users. Yunshan Jia, Chao Jin 0007, Qing Li 0028, Xuanzhe Liu, Xin Jin 0008 |
IEEE Trans. Netw. | 3 |
| 2025 | Niagara+: Scheduling Live ML Analytics Across Heterogeneous Device Processors and Edge ServersabstractIntelligent applications rely significantly on the live machine learning pipeline, a couple of deep neural network (DNN) inference services, executed on mobile devices to meet functional requirements while ensuring user data privacy. However, executing these DNN services on resource-constrained mobile devices presents a considerable challenge: low throughput and high energy consumption of inference tasks. To address this issue, we proposeNiagara+, a novel system designed to enhance throughput by jointly scheduling DNN inference services across heterogeneous processors on mobile devices and offloading services to powerful edge servers. To achieve this,Niagara+encounters two critical challenges: unpredictable workload dynamics and high scheduling complexity. To effectively tackle these challenges,Niagara+employs a predictive model to forecast incoming workload patterns and orchestrates service allocation across device heterogeneous processors and edge servers through a combination of two-step offline scheduling optimization and online service dispatching strategies. We implementedNiagara+and conducted comprehensive experiments, demonstrating its superiority over state-of-the-art approaches, reducing DNN service latency by up to 2.6× under high-bandwidth networks and 9.1× under low-bandwidth networks, while consistently meeting stringent inference latency requirements. Daliang Xu, Qing Li 0028, Mengwei Xu 0001, Gang Huang 0001, Shangguang Wang, Qun Wei, Xin Jin 0008, Yun Ma 0002, Xuanzhe Liu |
IEEE Trans. Serv. Comput. | 2 |
| 2025 | Flexible Shadow: Resource-Efficient Reliability Enhancement for Edge Services Through Dynamic Shadow CoordinationabstractEdge computing plays a pivotal role in supporting services necessitating sub-second latency, notably in domains like Industry 4.0 and autonomous driving. However, unpredictable failure occurring at edge servers can result in prolonged response time and decreased service reliability, posing significant risks to both safety and property. Traditional reliability mechanisms, namely task re-execution and task replication, are often inadequate for edge environments. The former struggles to meet the stringent end-to-end service latency requirements, while the latter imposes a high resource consumption burden on resource-limited edge clouds. To address this issue, this paper introduces a novel Flexible Shadow mechanism, where the backup instance, referred to as the Flexible Shadow, is allocated fewer computation resources compared to its primary instance to conserve computation resources, and temporally preempts a portion of resources from neighboring shadows to accelerate when necessary. To support the implementation of this mechanism, we propose the Flexible Shadow Backup Framework, a resource-efficient reliability enhancement framework for edge services through dynamic shadow coordination. This framework integrates three key components: a deployment algorithm for resource allocation, an adjustment algorithm for migration cost-latency tradeoffs, and a reconfiguration algorithm for adaptation optimization. Comprehensive experiments conducted on a Docker-based prototype demonstrate the effectiveness of the Flexible Shadow mechanism, achieving nearly 60% reduction in computing resource consumption compared to traditional approaches while maintaining sub-second latency. Lipei Yang, Ao Zhou 0001, Xiao Ma 0009, Qing Li 0028, Yuanzhe Li 0001, Shangguang Wang |
IEEE Trans. Serv. Comput. | 5 |
| 2024 | Flexible LAN-WAN Orchestration for Communication Efficient Federated Learning over Large-Scale Mobile DevicesabstractFederated learning (FL) has been widely adopted as a privacy-preserving model training paradigm. However, traditional FL protocol heavily relies on data transmission between clients and servers across the wide-area network (WAN), which is tightly constrained and unreliable, therefore causing expensive communication and slow convergence. To this end, we propose a LAN-aware FL (LanFL) protocol, which can efficiently leverage the network capacity of the local-area network (LAN). By frequent model aggregation among the devices within the same LAN, we can significantly reduce the global aggregation across WAN, thus accelerating the training process. However, due to the unique challenges introduced by LAN, it’s not easy to efficiently utilize LAN resources while preserving the original dignity of FL performance. Therefore, LanFL also incorporates several critical techniques: LAN-aware hierarchical aggregation, intraLAN device topology construction, and inter-LAN heterogeneous bandwidth coordination. Extensive real-world experiments are conducted and the experimental results show that LanFL can significantly accelerate FL training up to $6.0 \times$, while preserving the model accuracy. Jinliang Yuan, Qing Li 0028, Fan Dang 0001, Xiaofang Mu, Mengwei Xu 0001, Shangguang Wang |
ICPADS | 2 |
| 2024 | Deciphering the Enigma of Satellite Computing with COTS Devices: Measurement and AnalysisabstractIn the wake of the rapid deployment of large-scale low-Earth orbit satellite constellations, exploiting the full computing potential of Commercial Off-The-Shelf (COTS) devices in these environments has become a pressing issue. However, understanding this problem is far from straightforward due to the inherent differences between the terrestrial infrastructure and the satellite platform in space. In this paper, we take an important step towards closing this knowledge gap by presenting the first measurement study on the thermal control, power management, and performance of COTS computing devices on satellites. Our measurements reveal that the satellite platform and COTS computing devices significantly interplay in terms of the temperature and energy, forming the main constraints on satellite computing. Further, we analyze the critical factors that shape the characteristics of onboard COTS computing devices. We provide guidelines for future research on optimizing the use of such devices for computing purposes. Finally, we have released the datasets to facilitate further study in satellite computing. Ruolin Xing, Mengwei Xu 0001, Ao Zhou 0001, Qing Li 0028, Feng Qian 0001, Shangguang Wang |
MobiCom | 4 |
| 2024 | Exploring Real-Time Satellite Computing: From Energy and Thermal PerspectivesabstractSmall satellites (SmallSats) are now widely used in various fields, such as real-time communication and earth observation. These increasingly complex space applications face limited support from conventional radiation-hardened processors onboard. Hence, many SmallSats are designed to utilize high performance commercial off-the-shelf (COTS) computing devices to address this problem but it remains unclear how the unique energy and thermal characteristics of SmallSats impact computing efficiency onboard. This work conducts a systematic and quantitative measurement study of COTS devices’ computing efficiency on two real orbiting SmallSats. The key findings are: 1) inadequate energy management may lead to electricity wastage in sunlit zones and shortages in eclipse zones, impacting onboard computing availability and 2) the weak heat dissipation onboard may compromise COTS computing efficiency by incurring thermal throttling. To address such challenges, we design ProScale, a lightweight application-aware power management and thermal control system to improve computing efficiency under both electrical and thermal energy constraints. Evaluation shows that ProScale can improve the average task completion latency by $2.1 \times$ for computation-intensive applications compared with baselines. Qing Li 0028, Shangguang Wang, Chenren Xu, Xiao Ma 0009, Mengwei Xu 0001, Ao Zhou 0001, Ruolin Xing, Zuo Zhu, Ying Zhang 0012, Xuanzhe Liu |
RTSS | 1 |
| 2024 | Online Request Replication for Obtaining Fresh Information Under Pull ModelabstractAge of Information (AoI) has gained widespread usage and emerged as a pivotal metric for assessing timeliness performance in information-update systems. Such systems often entail service requirements for rapidly obtaining requested data in real-time. For instance, in the financial market, users rely on up-to-date and low-latency information to make appropriate trading decisions to maximize profits within their financial budge. Much of the existing research on real-time services focuses on ensuring AoI or service-level latency, but there is a growing demand for joint optimization of these two metrics to accommodate a broader range of potential applications. Therefore, this article investigates the problem of minimizing AoI within the context of statistical latency guarantees. To tackle the critical challenges posed by the joint modeling of AoI and statistical latency, system uncertainty, as well as tradeoff between performance and user’s budget, we employ a replication scheme to ensure both AoI and statistical service-level latency. To address the critical challenges posed by the unknown distribution in data updating processes and response times across providers, we formulate the AoI minimization problem with statistical latency constraints as a combinatorial multiarmed bandit problem utilizing the Lyapunov optimization theory. Subsequently, we propose an online learning-based request replication algorithm to address this problem. Our proposed algorithm achieves a cumulative regret of$O(T\sqrt {\log (T)})$compared to the genie-aided algorithm. Simulation results demonstrate the superior performance of the proposed algorithm against benchmarks. Qibo Sun, Qing Li 0028, Xiao Ma 0009, Shangguang Wang |
IEEE Internet Things J. | 3 |
| 2024 | Battery-Aware Energy Optimization for Satellite Edge ComputingabstractSatellite edge computing can incur dramatically increased energy demand onboard, which is met by satellite batteries during eclipses. Excessive energy usage during regular operations accelerates battery wear. Therefore, it is important and timely to optimize the energy consumption onboard to extend satellite batteries life. This paper investigates battery-aware energy optimization for satellite edge computing under energy harvesting dynamics and wireless environment uncertainty. Inspired by the periodical energy harvesting and satellite-ground connection, we develop a pattern-aware online energy scheduling algorithm within an online convex optimization framework. This learning algorithm achieves theoretical guarantees of no regret and gradually zeroing constraint violations. We further exploit inter-satellites collaboration to extend the average battery life in a whole constellation where satellites have different battery capacity degradation. Trace-driven simulations show that our algorithm can significantly extend the battery life by 1.32× and effectively adapt to the energy harvesting dynamics and wireless environment uncertainty. Qing Li 0028, Shangguang Wang, Xiao Ma 0009, Ao Zhou 0001, Yue Wang 0072, Gang Huang 0001, Xuanzhe Liu |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | Niagara: Scheduling DNN Inference Services on Heterogeneous Edge Processors
Daliang Xu, Qing Li 0028, Mengwei Xu 0001, Gang Huang 0001, Shangguang Wang, Xin Jin 0008, Yun Ma 0002, Xuanzhe Liu |
ICSOC (1) | 2 |
| 2023 | Satellite Computing: Vision and ChallengesabstractThe space industry experiences a rise in low-Earth-orbit satellite mega-constellations to achieve universal connectivity. At the same time, cloud firms (such as Google, Microsoft, and AWS) also have ambitions for computing in space to offer public cloud services in orbit. Satellite computing, as a new emerging concept, is promising to enable new paradigms in on-orbit autonomy, remote sensing, edge computing, and other areas by empowering satellites with computing resources. However, LEO mega-constellations bring inherent challenges for satellite computing in satellite networking, computing, and others due to the moving core infrastructure, reduced system power budget, and harsh space environment. This article presents vision and challenges for satellite computing based on a brief survey of the very recent literature in the “NewSpace” era and gives a case study of an open research platform on real satellites named Tiansuan constellation. This article aims to call researchers to collaboratively undertake the research of satellite computing and provide some insights for the research community. Shangguang Wang, Qing Li 0028 |
IEEE Internet Things J. | 2 |
| 2023 | Online Service Request Duplicating for Vehicular ApplicationsabstractVehicles on roads have increasingly powerful computing capabilities and edge nodes are being widely deployed. They can work together to provide computing services for onboard driving systems, passengers, and pedestrians. Typical applications in vehicular systems have service requirements such as low latency and high reliability. Most studies in vehicular networks concerning latency and reliability focus on vehicular communication at the network level. Based on these fundamental works, an increasing proportion of vehicles boast complex applications that require service-level end-to-end performance guarantees. Several works guarantee service-level latency or reliability while new and innovative applications are demanding a joint optimization of the above two metrics. To address the critical challenges induced by the joint modeling of latency and reliability, system uncertainty, and performance and cost trade-off, we employ service request duplication to ensure both latency and reliability performance at the service level. We propose an online learning-based service request duplication algorithm based on a multi-armed bandit framework and Lyapunov optimization theory. The proposed algorithm achieves an upper-bounded regret compared to the oracle algorithm. Simulations are based on real-world datasets and the results demonstrate that the proposed algorithm outperforms the benchmarks. Qing Li 0028, Xiao Ma 0009, Ao Zhou 0001, Changhee Joo, Shangguang Wang |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | User-Oriented Edge Node Grouping in Mobile Edge ComputingabstractIn mobile edge computing networks, densely deployed access points are empowered with computation and storage capacities. This brings benefits of enlarged edge capacity, ultra-low latency, and reduced backhaul congestion. This paper concerns edge node grouping in mobile edge computing, where multiple edge nodes serve one end user cooperatively to enhance user experience. Most existing studies focus on centralized schemes that have to collect global information and thus induce high overhead. Although some recent studies propose efficient decentralized schemes, most of them did not consider the system uncertainty from both the wireless environment and other users. To tackle the aforementioned problems, we first formulate the edge node grouping problem as a game that is proved to be an exact potential game with a unique Nash equilibrium. Then, we propose a novel decentralized learning-based edge node grouping algorithm, which guides users to make decisions by learning from historical feedback. Furthermore, we investigate two extended scenarios by generalizing our computation model and communication model, respectively. We further prove that our algorithms converge to the Nash equilibrium with upper-bounded learning loss. Simulation results show that our mechanisms can achieve up to 96.99% of the oracle benchmark. Qing Li 0028, Xiao Ma 0009, Ao Zhou 0001, Xiapu Luo, Fangchun Yang, Shangguang Wang |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | Dynamic Task Scheduling in Cloud-Assisted Mobile Edge ComputingabstractThe cloud-assisted mobile edge computing system is a critical architecture to process computation-intensive and delay-sensitive mobile applications in close proximity to mobile users with high resource efficiency. Due to the heterogenous dynamics of task arrivals at edge nodes and the distributed nature of the system, the workloads of edge nodes are prone to be unbalanced, which can cause high task response time and resource cost. This paper solves the dynamic task scheduling problem in cloud-assisted mobile edge computing (including both peer task scheduling among edge nodes and cross-layer task scheduling from edge nodes to the cloud), aiming at minimizing average task response time within resource budget limit. To overcome the challenges of task arrival dynamics, edge node heterogeneity, and computation-communication delay tradeoff, we propose aWater-filling BasedDynamic TaskScheduling (WiDaS) algorithm. WiDaS dynamically tunes the usage of cloud resources based on the Lyapunov optimization method and efficiently schedules mobile tasks among edge nodes (and the cloud) by exploiting the idea of water filling. Extensive simulations are conducted to evaluate WiDaS under a trace-driven traffic pattern and two mathematic traffic patterns. The results demonstrate that WiDaS shows two-fold benefits of efficiency and effectiveness. In terms of efficiency, WiDaS can achieve the approximate results with the KKT-based algorithm while reducing the computation complexity from exponential order to polynomial order. In terms of effectiveness, WiDaS can reduce the average task response time by up to 64.4% and 47.2% over the Fair-ratio and the Edge-first algorithm. Xiao Ma 0009, Ao Zhou 0001, Shan Zhang 0001, Qing Li 0028, Alex X. Liu, Shangguang Wang |
IEEE Trans. Mob. Comput. | 4 |
| 2022 | Service Coverage for Satellite Edge ComputingabstractRecently, increasing investments in satellite-related technologies make the low earth orbit (LEO) satellite constellation a strong complement to terrestrial networks. To mitigate the limitations of the traditional satellite constellation “bent-pipe” architecture, satellite edge computing (SEC) has been proposed by placing computing resources at the LEO satellite constellation. Most existing works focus on space-air-ground integrated network architecture and SEC computing framework. Beyond these works, we are the first to investigate how to efficiently deploy services on the SEC nodes to realize robustness aware service coverage with constrained resources. Facing the challenges of spatial-temporal system dynamics and service coverage-robustness conflict, we propose a novel online service placement algorithm with a theoretical performance guarantee by leveraging Lyapunov optimization and Gibbs sampling. Extensive simulation results show that our algorithm can improve the service coverage by$4.3\times $compared with the baseline. Qing Li 0028, Shangguang Wang, Xiao Ma 0009, Qibo Sun, Houpeng Wang, Suzhi Cao, Fangchun Yang |
IEEE Internet Things J. | 1 |
| 2022 | QoS Driven Task Offloading With Statistical Guarantee in Mobile Edge ComputingabstractIn mobile edge computing, popular mobile applications, such as augmented reality, usually offload their tasks to resource-rich edge servers. The user experience can be considerably affected when many mobile users compete for the limited communication and computation resources. The key technical challenge in task offloading is to guarantee the Quality of Service (QoS) for such applications. Existing work on task offloading focus on deterministic QoS (delay) guarantee, which means that tasks have to complete before the given deadline with 100 percent. However, it is impractical to impose a deterministic QoS guarantee for tasks due to the high dynamics of the wireless environment when offloading to edge servers. In this paper, we focus on task offloading with statistical QoS guarantee (tasks are allowed to complete before a given deadline with a probability above the given threshold), which can further save more energy by loosing the QoS requirement. Specially, we first propose a statistical computation model and a statistical transmission model to quantify the correlation between the statistical QoS guarantee and task offloading strategy. Then, we formulate the task offloading problem as a mixed integer non-Linear programming problem with the statistical delay constraint. We transform the statistical delay constraint into the constraints on CPU cycle numbers and the delay exponent respectively. We propose an algorithm to provide the statistical QoS guarantee for tasks using convex optimization theory and Gibbs sampling method. Experiment results show that the proposed algorithm outperforms the three baselines. Qing Li 0028, Shangguang Wang, Ao Zhou 0001, Xiao Ma 0009, Fangchun Yang, Alex X. Liu |
IEEE Trans. Mob. Comput. | 1 |