Pan Lai

dblp:139/7902 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0002-4967-5573ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 3 since 2021Computer networks · 6 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Multi-Modal Feature Fusion Distance Gating 3D Imaging Based on Edge Computing
abstract
3D Range-Gated Imaging technology is widely used for detection in complex environments (such as autonomous driving scenarios) due to its excellent anti-interference capabilities. However, its application faces the dual challenges of a lack of specialized datasets and the limited performance of traditional RGB models in low signal-to-noise ratio environments, which hinders the transfer and generalization of deep learning methods. To address these difficulties, this paper proposes a 3D imaging method based on multimodal feature fusion. Specifically, the model adopts a dual Vision Transformer (ViT) encoder, single-decoder architecture. On one hand, it performs pre-trained ViT encoding on geometrically re-projected RGB images. On the other hand, it applies an isomorphic ViT encoding to the range-gated images. Through layer-wise semantic recombination, it achieves efficient cross-modal feature fusion, not only does it enhance the robustness and accuracy of depth estimation, but it can also be easily deployed on edge devices. To overcome the problem of overfitting to LiDAR ground truth data, a spatially constrained window cropping data augmentation strategy is designed, significantly increasing the diversity of training samples and the model's generalization ability. To address the input resolution limitations of Transformers, an optimization scheme combining dynamic patch-based training and progressive up-sampling is further proposed, balancing high-resolution feature representation with efficient training. Experimental results show that the proposed method reduces the depth estimation RMSE on a public test set by more than 12% compared to mainstream baseline models, with particularly outstanding performance in low-texture and long-distance scenes. This research provides a systematic technical solution for cross-modal 3D perception and offers theoretical and engineering references for designing 3D imaging models for complex environments.
Yuanai Xie, Pan Lai, Xiao Zhang 0006, Jianlin Zhu
CloudCom4
2025 Communication-Efficient Scheduling for Cost- and Deadline-Aware Data Transfers in Support of Distributed Machine Learning
abstract
The success of distributed machine learning and federated learning in cloud environments depends critically on efficient communication across heterogeneous and bandwidth-constrained networks. While the exchange of raw data is often avoided due to privacy and regulatory concerns, many distributed workflows still require large-scale data movement or model updates that demand predictable network performance. This paper investigates the trade-off between earliest completion time (ECT) and communication cost in data transfers that leverage bandwidth reservation on fixed paths with variable bandwidth in dedicated high-performance networks. We fo-cus on two representative scheduling objectives relevant to distributed learning and data-intensive cloud workflows: (i) minimizing communication cost while meeting a transfer dead-line, and (ii) achieving the ECT while respecting a maximum cost budget. We prove that both scheduling problems are NP-complete and propose heuristic algorithms to address them effectively. Extensive simulations demonstrate that our methods can substantially reduce both latency and cost, providing a prac-tical foundation for communication-efficient and resource-aware distributed machine learning in cloud computing environments.
Liudong Zuo, Pan Lai
CloudCom2
2025 Intelligent Autoscaling of Microservice and Request Routing for Dynamic Service Requests
abstract
Microservice architecture provides innovative solutions for delay-sensitive applications and is widely used in Mobile Edge Computing (MEC). Deploying microservices and implementing request routing within large-scale networks are confronted with numerous challenges, primarily due to intricate dependencies between microservices, frequent data communications between microservices, and the dynamic fluctuations of user request traffic. Given that the user requests are time-varying, the orchestration scheme must enable automatic scaling of microservice instances to ensure service quality. Yet, existing research mainly focuses on static microservice instance deployment, inadequately achieving the intelligent autoscaling of microservices and addressing the time-varying nature of user requests. To address these challenges in dynamic MEC networks, this paper introduces a joint optimization strategy for microservice autoscaling and routing, which adapts to the dynamic fluctuations of user request traffic. Initially, we utilize open Jackson queuing network theory to construct the model, analyzing service request queuing, communication, and processing delays along routing paths, and formulate the problem aiming to minimize deployment costs while ensuring service processing delays are not beyond the delay range that the users can accept. To address the problem, we propose a multi-stage, fine-grained dynamic scaling and routing algorithm. Extensive simulation results indicate that our approach substantially reduces network costs while maintaining network delays within a reasonable range, compared to the other state-of-the-art methods.
Pan Lai, Yang Chen 0072, Shisheng Lin, Tongxin Liao, Menglan Hu, Xiao Zhang 0006, Yuanai Xie
IEEE Internet Things J.1
2025 Reflection Optimization for Covert Ambient Backscatter Systems Under Two Jamming Patterns
abstract
Ambient backscatter communication (ABC) enables low-cost and energy-efficient connectivity for Internet of Things (IoT) devices by leveraging ambient radio-frequency (RF) signals. However, the passive nature and open wireless medium of ABC systems make them vulnerable to detection by unauthorized receivers (wardens). To mitigate this risk, covert communication, which conceals transmissions by embedding them within noise, offers a promising security enhancement for ABC systems. This paper proposes a jammer-assisted reflection coefficient optimization framework to enhance the covertness and reliability of ABC systems with an endogenous warden and an external jammer. Specifically, we consider two distinct jamming patterns: uniformly distributed and truncated exponentially distributed artificial noise power. We derive closed-form expressions for both the outage probability of the backscatter link and the minimum detection error rate at the warden under these jamming patterns. Based on these expressions, we determine the optimal reflection coefficients that maximize the effective covert rate while satisfying a predefined covertness constraint. Additionally, we introduce the concept of jamming cost to evaluate the efficiency and applicability of different jamming patterns in terms of the required jamming power to achieve a desired level of covertness. Numerical results validate the effectiveness of the proposed optimization framework and reveal that while uniform jamming provides stronger covertness and lower jamming cost, truncated exponential jamming achieves a lower outage probability. These findings provide key insights for designing secure and efficient ABC systems across diverse IoT deployment scenarios.
Yuanai Xie, Yaoyao Wen, Xiao Zhang 0006, Pan Lai, Zhixin Liu 0001, Haoyuan Pan, Tse-Tin Chan
IEEE Internet Things J.4
2024 Optimal Scheduling Algorithms for Cost-Effective Bandwidth Reservation in HPNs
abstract
Vast amounts of data are continually being produced in various scientific fields. Once these large datasets are generated, they often require rapid transfer over long distances using bandwidth reservation services provided by high-performance networks (HPNs) dedicated to collaborative data storage and analysis. The primary goal of data transfer is typically to achieve the earliest completion time (ECT), but users may also seek to minimize financial costs associated with the transfer. Balancing these differing requirements can be challenging. In this paper, we explore the trade-off between ECT and cost in data transfers that use bandwidth reservation on variable paths with fixed bandwidth within dedicated HPNs. Our investigation focuses on two types of bandwidth reservation requests (BRRs) and their scheduling: (i) minimizing data transfer cost while meeting a data transfer deadline, and (ii) achieving ECT while meeting a specified maximum cost. To optimize the scheduling of both types of BRRs, we propose two novel algorithms and conduct extensive simulations to demonstrate their effectiveness and efficiency.
Liudong Zuo, Pan Lai, Zhong Chen 0003
IEEE Big Data2
2024 Reinforcement Learning for Efficient Multi-phase Resource Allocation
abstract
Efficient resource allocation is pivotal for achieving high performance in emerging computer systems, where multiple users and tasks compete for shared resources. This challenge spans various domains, including data centers, multicore processors, cloud computing and edge computing, each requiring nuanced allocation strategies to balance competing demands. Traditional approaches often assume concave utility (performance) functions for users, simplifying optimization but failing to capture the complexities of real-world scenarios where non-concave utility functions prevail. Numerous works in the literature apply the greedy algorithm to nonconcave utility functions, resulting in suboptimal solution due to the short-sighted behaviors. To improve this gap, we propose a novel multiphase resource allocation framework that accurately reflects the non-linear dynamics of these systems. To tackle the NP-complete nature of this problem, we formulate a customized resource allocation Markov Decision Process (MDP) that integrates the characteristics of multi-phase utility functions into a nuanced design of the key MDP components, such as state representations, reward signals, and actions. We explore two reinforcement learning (RL)-based methods, specifically Dueling Deep Q-Network (Dueling DQN) and Proximal Policy Optimization (PPO), to optimize resource allocation over time. Our RL-based strategies outperform the conventional greedy algorithm by approximately 37% in standard environments and up to 73% in specialized environments, highlighting their effectiveness in handling the resource allocation problem with non-concave utility functions and achieving scalable, real-time solutions.
Zhenfu Zhang, Haiyan Yin, Liudong Zuo, Xiao Zhang 0006, Jianlin Zhu, Yuxuan Fan, Pan Lai
HPCC7
2024 Online Dynamic Scaling of Microservices with Fair Probabilistic Routing
abstract
Microservice architecture, as an emerging network architecture, has gained widespread adoption in latency-sensitive applications within the realm of mobile edge computing (MEC). In MEC networks, these latency-sensitive applications necessitate the concurrent processing of numerous service requests, which are composed of microservices. The complex dependencies between microservices and frequent data communication between servers contribute to the intricacy of deploying and routing microservice instances within the network. Moreover, the dynamic and unpredictable nature of service request traffic significantly complicates the timeliness and efficiency of service deployment and request routing strategies. However, existing research predominantly focuses on static network environments and neglects the time-varying characteristics of service request traffic in realistic scenarios. Consequently, we address the joint optimization problem of service deployment and request routing in the presence of dynamic service request traffic. To model the inherent data dependencies and analyze service request response latency, we employ the open Jackson queuing network. We propose a fine-grained microservice dynamic scaling (FMDS) algorithm to capture the dynamic fluctuations in service request traffic within the network. This algorithm scales microservice instances based on the principle of equal proportional change, obtaining a service deployment scheme that minimizes costs while satisfying latency constraints. Furthermore, we introduce a recursive path search algorithm that explores the service deployment scheme to determine the node forwarding probability for the entire network, adhering to the principles of fair routing. Simulation results show that the proposed method effectively improves network latency stability by 75% and enhances the timeliness of the service deployment strategy.
Yang Chen 0072, Shisheng Lin, Liangyuan Wang, Menglan Hu, Pan Lai, Yuanai Xie
ISPA6
2024 Dynamic Task Scheduling for Coordinated Truck-Drone Parcel Delivery: A Hybrid Genetic Tabu Algorithm Approach
abstract
With the rapid growth of the on-demand economy, logistics companies and merchants increasingly struggle to meet customer demands in dynamic and uncertain conditions. This paper studies the coordinated delivery of parcels by trucks and drones under such demands, proposing a Dynamic Task Scheduling Algorithm based on Hybrid Genetic Tabu algorithm (DTSAGT) for route optimization. Simulating dynamic customer demands with a Poisson distribution and statistical methods, the algorithm addresses timeliness issues due to variations in customer needs. It optimizes drone path planning and task allocation considering drone endurance and payload limits to minimize total delivery time. The algorithm includes three steps: initial solution construction, iterative optimization, and dynamic operations. Experimental results show that DTSAGT reduces the total service time by 15.05%, 34.13%, and 35.71% on average compared to the baseline algorithms. This paper’s contribution is the combination of hybrid genetic and tabu search algorithms applied to dynamic task scheduling in truck-drone delivery, enhancing logistics efficiency.
Ruitai Li, Tongxin Liao, Lijun Luo, Menglan Hu, Pan Lai, Xiao Zhang 0006
ISPA6
2024 Joint optimization of application placement and resource allocation for enhanced performance in heterogeneous multi-server systems
Pan Lai, Yiran Tao, Yuanai Xie, Shanjiang Tang, Shengquan Liao
Comput. Networks1
2023 Reflection-Optimized Covert Communication for Jammer-Aided Ambient Backscatter Systems
abstract
The integration of Ambient Backscatter Communication (ABC) with covert communication is expected to support emerging Internet of Things (IoT) applications (e.g., Radio Frequency (RF)-powered networks) due to the need for low-cost connectivity and confidential transmission. In general, the purpose of covert communication is to hide the existence of the RF-powered wireless link to ensure the information security of the ABC link. However, the ABC link may have a high rate requirement, thus inevitably increasing the risk of information leakage. Hence, this paper considers jammer-aided endogenous covert communication, where an RF tag sends information covertly to an ABC receiver and exploits the jammer's Artificial Noise (AN) under the supervision of a warden-like legacy receiver. To obtain the maximum data rate of the backscatter link without being detected, we derive the minimum detection error rate of the warden and the outage probability of the backscatter link under random channel fading and the jammer's AN, respectively. Then, we optimize the tag's reflection coefficient to maximize its effective covert rate under the covert constraint based on the warden's mean detection error rate. Since the optimal reflection coefficient cannot be solved directly, monotonicity analyses of the objective and the constraint with respect to the reflection coefficient are adopted to achieve an efficient solution. Numerical results demonstrate the effectiveness of the optimized reflection coefficient for the jammer-aided system.
Yuanai Xie, Tse-Tin Chan, Xiao Zhang 0006, Pan Lai, Haoyuan Pan
GLOBECOM4
2022 Dynamic thresholding for video anomaly detection
abstract
Abstract Anomaly detection is one of the most important applications in video surveillance that involves the temporal localisation of anomaly events in unannotated video sequences. By learning the normal patterns to generate frames and calculating their reconstruction error relative to the ground truth, a frame can be recognised as being abnormal if the reconstruction error exceeds a threshold. Most existing works use a fixed threshold that computes over all the testing data to determine the anomalies. However, fixed threshold strategy cannot address the challenges brought by the dynamic environment, e.g. changes in illumination conditions. In this paper, a dynamic thresholding algorithm (DTA) is proposed, which is fully data‐driven and capable of automatically determining thresholds such that the developed anomaly detection system can flexibly adapt to different scenarios. The proposed DTA is independent of the backbone network and can be easily incorporated into most existing video anomaly detection models to help identify the appropriate thresholds. On both synthetic and real‐world datasets, the experimental results show that with the proposed DTA, the video anomaly detection methods achieve a better performance considering the changes in dynamic environment.
Diyang Jia, Xiao Zhang 0006, Joey Tianyi Zhou, Pan Lai, Yifei Wei
IET Image Process.4
2022 Utility Optimal Thread Assignment and Resource Allocation in Multi-Server Systems
abstract
Achieving high performance in many multi-server systems (e.g., web hosting center, cloud) requires finding a good assignment of worker threads to servers and also effectively allocating each server’s resources to its assigned threads. The assignment and allocation components of this problem have been studied extensively but largely separately in the literature. In this paper, we introduce theassign and allocate (AA)problem, which seeks to simultaneously find an assignment and allocation that maximizes the total utility of the threads. Assigning and allocating the threads together can result in substantially better overall utility than performing the steps separately, as is traditionally done. We model each thread by a utility function giving its performance as a function of its assigned resources. We first prove that the AA problem is NP-hard. We then present a$2 (\sqrt {2}-1) > 0.828$factor approximation algorithm for concave utility functions, which runs in$O(mn^{2} + n (\log mC)^{2})$time for$n$threads and$m$servers with$C$amount of resources each. We also give a faster algorithm with the same approximation ratio and$O(n (\log mC)^{2})$time complexity. We then extend the problem to two more general settings. First, we consider threads with nonconcave utility functions, and give a 1/2 factor approximation algorithm. Next, we give an algorithm for threads using multiple types of resources, and show the algorithm achieves good empirical performance. We conduct extensive experiments to test the performance of our algorithms on threads with both synthetic and realistic utility functions, and find that they achieve over 92% of the optimal utility on average. We also compare our algorithms with a number of practical heuristics, and find that our algorithms achieve up to 9 times higher total utility.
Pan Lai, Rui Fan 0004, Xiao Zhang 0006, Wei Zhang 0082, Fang Liu 0009, Joey Tianyi Zhou
IEEE/ACM Trans. Netw.1
2016 Utility Maximizing Thread Assignment and Resource Allocation
abstract
Achieving high performance in many distributed systems requires finding a good assignment of threads to servers as well as effectively allocating each server's resources to its assigned threads. The assignment and allocation components of this problem have both been studied extensively, but separately in the literature. In this paper, we introduce the assign and allocate (AA) problem, which seeks to simultaneously find an assignment and allocations that maximize the total utility of the threads. Assigning and allocating the threads together can result in substantially better overall utility than performing the steps separately, as is traditionally done. We model each thread by a concave utility function giving its throughput as a function of its assigned resources. We first show that the AA problem is NP-hard, even when there are only two servers. We then present a 2(√2-1) > 0.828 factor approximation algorithm, which runs in O(mn2 + n (log mC)2) time for n threads and m servers with C amount of resources each. We also present a faster algorithm with the same approximation ratio and O(n(log mC)2) running time. We conducted experiments to test the performance of our algorithm on threads with different types of utility functions, and found that it achieves over 99% of the optimal utility on average. We also compared our algorithm against several other assignment and allocation algorithms, and found that it achieves up to 5.7 times better total utility.
Pan Lai, Rui Fan 0004, Wei Zhang 0082, Fang Liu 0009
IPDPS1
2015 Energy-Aware Caching
abstract
To achieve higher performance, cache sizes have been steadily increasing in computer processors and network systems. But caches are often over-provisioned for peak demand and underutilized in typical non-peak workloads. As caches consume substantial power, this results in significant amounts of wasted energy. To address this, existing works turn off parts of the cache when they do not contribute to higher performance. However, while these methods are effective empirically, they lack provable performance bounds. In addition, existing works focus on processor caches and are not applicable to network caches where data size and cost can vary. In this paper, we study the energy-aware caching (EAC) problem, and seek to minimize the total cost incurred due to cache misses and energy consumption. We propose three algorithms to solve different variants of this problem. The first is an optimal offline algorithm that runs in O(kn log n) time for a size k cache and n cache accesses. Then, we propose a simple online algorithm for uniform data size and cost that is $2 + {{h} \over {h-h+1}}$ competitive compared to an optimal algorithm with a size h ≤ k cache. Lastly, we propose a $2 + {{h-1} \over {h-h+1}}$ competitive online algorithm that allows arbitrary data sizes and costs. We give an efficient implementation of the algorithm that takes O(log k) amortized time per cache access, and also present an adaptive version that reacts to workload patterns to achieve better real-world performance. Using trace driven simulations, we show our algorithm has substantially lower cost than algorithms focused on maximizing cache hit rates or minimizing energy usage alone.
Wei Zhang 0082, Rui Fan 0004, Fang Liu 0009, Pan Lai
ICPADS4
2015 Fast optimal nonconcave resource allocation
abstract
Efficient use of shared resources is a key problem in a wide range of computer systems, from cloud computing to multicore processors. Optimized allocation of resources among users can result in dramatically improved overall system performance. Resource allocation is in general NP-complete, and past works have mostly focused on studying concave performance curves, applying heuristics to nonconcave curves, or finding optimal solutions using slow dynamic programming methods. These approaches have drawbacks in terms of generality, accuracy and efficiency. In this paper, we observe that realistic performance curves are often not concave, but rather can be broken into a small number of concave or convex segments. We present efficient algorithms for optimal and approximately optimal resource allocation leveraging this idea. We also introduce several algorithmic techniques that may be of independent interest. Our optimal algorithm runs in O(snα(m)m(log m)2) time, and our approximation algorithm finds a 1 - ε optimal allocation for any ε > 0 in O(s/ε α(n/ε)n2log n/ε log m) time; here, s is the number of segments, n the number of processes, m the amount of shared resource, and α is the inverse Ackermann function that is ≤ 4 in practice. Existing exact and approximation algorithms have O(nm2) and O(n2m/ε) running times, resp., so our algorithms are much faster in the practical case where n <;<; m. Experiments show that our algorithms are 215 times faster than dynamic programming for finding optimal solutions when m = 1M, and produce solutions with 33% better performance than greedy algorithms.
Pan Lai
INFOCOM1
2014 A secure routing model based on distance vector routing algorithm
Bin Wang 0062, Chunming Wu 0001, Qiang Yang 0004, Pan Lai, Julong Lan
Sci. China Inf. Sci.4
2013 Makespan-Optimal Cache Partitioning
abstract
In current multicore systems, cache memory is shared between multiple concurrent threads. Allocating the proper amount of cache to each thread is crucial to achieving high performance. Cache management in many existing systems is based on the least recently used replacement policy, which can lead to adverse contention between threads for shared cache space. Cache partitioning is a technique that reserves a certain amount of cache for each thread, and has been shown to work well in practice. We introduce the problem of determining the optimal cache partitioning to minimize the make span for completing a set of tasks. We analyze the problem using a model that generalizes a widely used empirical model for cache miss rates. Our first contribution is to give a mathematical characterization of the properties satisfied by an optimal partitioning. Second, we present an algorithm that finds a 1 +\epsilon approximation to the optimal partitioning in O(n log \frac{n}{\epsilon}log\frac{n}{\epsilon p}) time, where n is the number of tasks and p is a value that depends on the optimal solution. We compare our algorithm with several partitioning schemes used in practice or proposed in the literature. Simulations show that our algorithm achieves between 22-59% better make span compared to these algorithms.
Pan Lai
MASCOTS1