Long Chen 0021

dblp:64/5725-21 · DBLP profile ↗
← Back
27ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0003-0543-7060ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 4 first-author · 5 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorComputer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DyMerge-LoRA: On-GPU Post-Merge Fusion for High-Throughput Multi-Tenant Composite LoRA Serving
abstract
In this paper, we consider the problem of efficiently serving multi-tenant requests for large language models (LLMs) equipped with compositions of multiple Low-Rank Adaptation (LoRA) adapters, aiming to optimize inference throughput, latency, and GPU memory usage. Existing multi-tenant inference systems are primarily architected for a one-to-one mapping between requests and single adapters, limiting their efficiency in handling compositions of multiple LoRA adapters. To address these issues, we propose DyMerge-LoRA, a novel inference serving system specifically designed for high-throughput multi-tenant scenarios by dynamically fusing corresponding LoRA adapters on the GPU. DyMerge-LoRA introduces a Base Adapter Manager (BAM) for efficient GPU memory management of individual basic adapters, a LoRA-Aware Greedy Scheduler (LAG) for batching, and a Mapping-based Adapter Fusion Matrix-Vector Multiplication (MAF-MVM) kernel for high computational throughput. Experimental evaluations conducted on multiple setups demonstrate that DyMerge-LoRA achieves up to 1.1-4× improvement in inference throughput compared to state-of-the-art methods, significantly reduces TTFT, and effectively mitigates GPU memory consumption, validating its effectiveness and performance advantages under diverse workload conditions.
Long Chen 0021, Huazheng Lao
KDD (1)2
2026 Reliable Truth Discovery for Dynamic and Dependent Sources
abstract
In the era of Big Data and generative artificial intelligence (AI), discovering the truth about various objects from different sources has become a pressing topic. Existing studies primarily focus on dependent sources with conflicting information, where sources may copy information from each other. However, real-world scenarios are often more complex, with dynamic dependence relationships among sources over time. This complexity makes it much more difficult to discover the truth. One of the key challenges centers on measuring the dynamic dependence among sources. To address this challenge, we have developed three models:$Depen\_{S}imple$,$Depen\_{C}omplex$, and$Depen\_{D}ynamic$. These models are based on the Hidden Markov Model (HMM) and are designed to handle different types of dependencies, namelysimple source dependence,complex source dependence, anddynamic source dependence. Based on the constructed models, we propose a generic framework for discovering the latent truth which are evaluated by three HMM-based methods. We conduct extensive experiments on three real-world datasets to evaluate the performance of the proposed methods, and the results demonstrate that all three methods achieve high accuracy over the state-of-the-art methods.
He Zhang 0028, Shuang Wang 0012, Long Chen 0021, Xiaoping Li 0001, Qing Gao 0001, Quan Z. Sheng
IEEE Trans. Knowl. Data Eng.3
2026 Container-State-Aware Energy Consumption Minimization for Serverless Functions
abstract
In this paper, we investigate the problem of minimizing the total energy consumption for serverless functions (SFs) while ensuring that the cold start probability of each SF remains below a specified threshold. In serverless computing, containers are reused to reduce the occurrence of cold starts, which increases the static energy consumption and decreases the dynamic energy consumption of physical machines (PMs). The trade-off between cold start reduction and energy consumption introduces significant challenges, especially considering that container states (idle, cold start, and running) have distinct energy characteristics. To address these challenges, we propose a container-state-aware energy consumption model that accurately captures the dynamic energy characteristics of PMs. Furthermore, we develop an enhanced Bayesian Optimization-Based Energy Consumption Minimization (BOECM) algorithm, which determines the optimal reuse times for different serverless functions (SFs) by decomposing the high-dimensional search space into a series of lower-dimensional sub-problems. In addition, several tailored heuristic strategies are proposed for request allocation and container deployment. The proposed BOECM algorithm is compared with several existing heuristic and meta-heuristic algorithms for similar problems. Experimental results demonstrate that the proposed method significantly reduces total energy consumption, maintains cold start probabilities within acceptable limits, and outperforms the baselines.
Long Chen 0021, Xiaoping Li 0001, Bing Ai
IEEE Trans. Serv. Comput.3
2026 DDRL: A Dual-Phase Deep Reinforcement Learning Approach for UAV-Assisted Content Delivery Across Multiple Base Stations
abstract
Uncrewed Aerial Vehicles (UAVs) with caching capabilities present a flexible and scalable approach for efficient content delivery in wireless communication environments with high demand. However, the challenges posed by the limited energy and storage capabilities of UAVs significantly affect content delivery efficiency. In this paper, we consider the problem of UAV-assisted content delivery across multiple base stations, with the aim of reducing content acquisition delays by jointly optimizing UAV trajectory, cache replacement, and transmission power. A Dual-phase Deep Reinforcement Learning (DDRL) framework is proposed, integrating real-time decision making with offline training to adapt dynamic user demands and multi-BS configurations. The Particle Swarm Optimization (PSO) algorithm is incorporated to improve UAV caching performance. The simulation results demonstrate that the DDRL framework achieves up to a 8% reduction in latency and a 5% improvement in cache hit rate compared to the best baseline algorithm, showcasing its efficiency in UAV-assisted content delivery.
Xinshuai Hua, Long Chen 0021, Xiaoping Li 0001
IEEE Trans. Wirel. Commun.2
2025 CAFGD: Reduce Fragmentation in Large-Scale Multi-tenant Clusters for GPU Sharing Workloads
abstract
GPU fragmentation has emerged as a critical obstacle to maximizing resource utilization in large-scale multi-tenant GPU clusters, despite advancements in GPU sharing techniques. This issue stems from the tidal nature of load fluctuations, the hierarchy of resource allocation, and variations in task runtime and resource requests. To address these challenges, we propose the Cluster-Aware Fragmentation Gradient Descent (CAFGD) algorithm, which formulates resource allocation as a hierarchical multi-dimensional packing problem with time dimension. CAFGD dynamically evaluates GPU fragmentation by considering task life cycles, balances CPU and GPU resources through a counterweight approach, and optimizes the allocation structure during low-demand periods with an optimistic waiting strategy aligned to cluster periodicity. Experiments on real-world traces with over 50,000 tasks demonstrate that CAFGD reduces unallocated GPU resources by more than 13.93% compared to existing packing-based schedulers, showcasing its effectiveness in enhancing cluster efficiency and GPU utilization.
Huazheng Lao, Long Chen 0021
ICDCS3
2025 Adaptive density-based clustering for many objective similarity or redundancy evolutionary optimization
Mingjing Wang, Ali Asghar Heidari, Long Chen 0021, Ruili Wang 0001, Mingzhe Liu 0001, Lizhi Shao, Huiling Chen 0001
Expert Syst. Appl.3
2025 ServlessSimPro: A comprehensive serverless simulation platform
Long Chen 0021
Future Gener. Comput. Syst.3
2025 AIHO: Enhancing task offloading and reducing latency in serverless multi-edge-to-cloud systems
Xin Li 0124, Long Chen 0021, Zian Yuan, Guangrui Liu
Future Gener. Comput. Syst.2
2024 Reinforcement Learning Based Memory Configuration for Linear Dynamic Function Chains
abstract
Serverless applications based on microservices typically comprise dozens or hundreds of loosely coupled functions. Each request in these applications triggers a function chain, which invokes a subset of functions. However, the invocation paths may not be predetermined in dynamic function chains. In addition, there are multiple memory configurations for each function in a chain. Different memory configuration combinations bring different execution times and costs. The configuration space grows exponentially as the length of the function chain increases. Therefore, it is challenging to determine the memory configurations for functions in a dynamic function chain to minimize the execution time and cost with an uncertain invocation path and a large configuration space. In this paper, a memory configuration problem of linear dynamic function chains is investigated to minimize the cost while meeting a specified SLO (service level objective). RLMC (Reinforcement Learning-based Memory Configuration Algorithm) is adopted to make memory configuration decisions dynamically. States, actions, and rewards are specially designed for the problem under study. In addition, a punishment factor adjustment strategy is developed to accommodate different SLOs. The proposed algorithm is evaluated and compared to existing algorithms over a comprehensive set of randomly generated serverless workflow applications. Experimental results demonstrate that RLMC significantly reduces the cost of dynamic function chains while meeting SLOs and outperforms other algorithms.
Xiaoping Li 0001, Kamran Yaseen Rajput, Long Chen 0021
CSCWD4
2024 HMUR: A Two-Stage Heuristic for UAV Scheduling in Mobile Edge Computing with Time Window Constraints
abstract
In Mobile Edge Computing (MEC), efficiently scheduling limited resources to fulfill the service requirements of massive Internet of Things (IoT) ground devices (GDs) presents a significant challenge. Unmanned Aerial Vehicles (UAVs), as airborne mobile devices, can be deployed to GD Task Areas (TAs) to provide communication relay or computational support. In this paper, a new multiple UAVs scheduling problem with UAV heterogeneity and time windows constraints is considered and modeled as a Multi-Traveling Salesman Problem (MTSP) with soft time constraints. A two-stage heuristic Heterogeneous Multiple UAVs Routing (HMUR) was proposed for the problem. The approach first identifies TAs and the optimal hovering positions for UAVs, and defines an effective fitness measurement to match between UAVs and different routing path under certain energy constraints. A score function is also defined to determine the final path under the real-time constraints of tasks, thereby enhancing the Quality of Service (QoS). Simulation results demonstrate that our proposed HMUR method surpasses existing baseline algorithms across multiple metrics, validating its effectiveness in optimizing resource scheduling in MEC environments.
Guangrui Liu, Long Chen 0021, Xin Li 0124
ICWS3
2024 Scheduling Workflows With Limited Budget to Cloud Server and Serverless Resources
abstract
Serverless functions (SFs) and on-demand virtual machines (VMs) are common cloud resources for scientific workflow applications, which are widespread in many fields. SFs are paid by actual running time with higher unit costs and higher resource utilization than VMs which are paid by billing time units. Generally, each application is executed on a limited budget. In this article, we study the challenging cloud workflow scheduling problem with a limited budget to minimize makespan in a hybridization of SFs and on-demand VMs for which the BCWS (Budget Constrained Workflow Scheduling) algorithm is proposed. Methods are developed to determine the task execution order, rent cloud resources and map tasks to resources respectively. Together with initial schedule construction and schedule improvement policies, these procedures are repeatedly applied in BCWS. The proposed algorithm is evaluated by comparing it to existing algorithms for similar problems over a comprehensive set of workflow instances. Experimental results show that the proposed algorithm significantly reduces the makespan with a hybrid configuration of VMs and SFs compared to the server only or the serverless only configurations and outperforms the compared algorithms which are the best existing ones for similar problems.
Xiaoping Li 0001, Long Chen 0021, Rubén Ruiz
IEEE Trans. Serv. Comput.3
2023 Smart Offloading Computation-intensive & Delay-intensive Tasks of Real-time Workflows in Mobile Edge Computing
abstract
In MEC, many deadline-constrained real-time work-flows with computation-intensive and/or delay-sensitive tasks are common in intelligent mobile devices (MDs). Though a task can be executed by either the local MD or an MEC server, the tasks of each work-flow are constrained by complex precedences, and real-time task offloading is somewhat tricky. In this paper, we consider the task offloading problem for stochastic work-flows with soft deadline constraints to minimize total tardiness and proposed an online RL-based offloading algorithm. In the algorithm, realtime tasks are dynamically partitioned into partial precedences in terms of which real-time RL states are constructed. Adaptive offloading actions are developed to determine task execution sequences for different states to optimize total tardiness. Experimental results show that the proposed online offloading algorithm outperforms the compared ones.
Haihong Zhu, Xiaoping Li 0001, Long Chen 0021, Rubén Ruiz
ICWS3
2023 Medical machine learning based on multiobjective evolutionary algorithm using learning decomposition
Mingjing Wang, Xiaoping Li 0001, Long Chen 0021, Huiling Chen 0001
Expert Syst. Appl.3
2023 An incremental learning evolutionary algorithm for many-objective optimization with irregular Pareto fronts
Mingjing Wang, Xiaoping Li 0001, Long Chen 0021, Huiling Chen 0001, Rubén Ruiz
Inf. Sci.4
2022 A Bi-Objective Learn-and-Deploy Scheduling Method for Bursty and Stochastic Requests on Heterogeneous Cloud Servers
abstract
In this article, we consider the dynamic allocation of bursty requests stochastically arriving at heterogeneous servers with uncertain setup times. Lower expected response time and less power consumption are desirable objectives of users and service providers respectively. However, sudden increase and decrease of cloud servers caused by bursty requests are rather challenging to get an appropriate trade-off between the two conflicting objectives which are closely related to the launched servers. The heterogeneity of the cloud servers further makes it more difficult to decide how to switch on and off servers and effectively and efficiently allocate bursty requests with balanced objectives. Based on a Markov decision process, a real-time bilevel decision-making model is constructed for unallocated requests which includes: whether to launch a server and which type of server to launch. A learn-and-deploy algorithm framework is proposed which contains two complementary stages. In the first stage, an effective offline bi-objective optimization algorithm is proposed to learn a set of policies, which provides helpful trade-off information for a decision-maker to choose a preferred policya posteriori. In terms of the system status, a policy decides whether to launch a server according to a state-action table and which server to launch using a server priority sequence. In the second stage, a computationally efficient policy deployment method is proposed to search the corresponding action in the selected policy based on the current system status and apply it to the real-time system. Experimental studies over a large number of random and real instances have been conducted to validate the effectiveness of the proposed bilevel model and algorithm. Compared to the most recent existing method, the performance of the proposed approach can at most achieve an 80% improvement on power consumption and 20% improvement on response time.
Xinye Cai, Xiaoping Li 0001, Long Chen 0021, Rubén Ruiz García, Qingfu Zhang 0001
IEEE Trans. Parallel Distributed Syst.5
2021 Multi-Tenant Cloud-Edge Workflow Scheduling With Priority and Deadline Constraints
abstract
The maximization of the Quality of Service (QoS) for multi-tenants is one of the key issues for cloud-edge service providers. Limited computing resources, different priorities, and deadlines of tenants make it difficult to satisfy the demands of all the multi-tenants. This paper considers the problem of scheduling limited cloud-edge resources to multi-tenant workflow applications with priority and deadline constraints. A level-based iterative greedy algorithm for the problem is proposed. The algorithm defines a priority-based multi-tenant instance success entropy to measure the total quality of service. The destruction & reconstruction and local search of the algorithm is performed based on the level of tasks. The proposed algorithm is compared to modified classical algorithms for similar problems. Experimental results demonstrate the effectiveness of the proposal for the considered problem.
Dongyuan Pan, Long Chen 0021, Xiaoping Li 0001
SERVICES2
2021 Hybrid Cloud Resource Scheduling With Multi-dimensional Configuration Requirements
abstract
Task scheduling with multi-dimensional configuration requirements is widely used in cloud platforms such as OpenStack and Kubernetes. In this paper, we consider the problem of scheduling tasks with multi-dimensional configuration to hybrid resources. An energy-aware scheduling algorithm on tasks with multi-dimensional configuration requirements (ESMCR in short) is presented. ESMCR is combined with a decomposition-based multi-objective evolutionary algorithm to minimize energy consumption and provide sufficient capacity for the data center. An entropy-based performance index is modeled to measure the QoS. The experimental results indicate that the proposed algorithms outperform the compared algorithms significantly.
Zhaokun Qiu, Long Chen 0021, Xiaoping Li 0001
SERVICES2
2021 Hybrid Resource Provisioning for Cloud Workflows with Malleable and Rigid Tasks
abstract
In cloud computing, reserved and on-demand instances are generally provided by service providers. Hybridization of the two alternatives can considerably save costs when renting resources from the cloud. However, it is a big challenge to determine the appropriate amount of reserved and on-demand resources in terms of users’ requirements. In this paper, the workflow scheduling problem with both reserved and on-demand instances is considered. The objective is to minimize the total rental cost under deadline constrains. The considered problem is mathematically modeled. A multiple sequence-based earliest finish time method is proposed to construct schedules for the workflows. Four different rules are used to generate initial task allocation sequences. Types and quantities of resources are determined by a free time block-based schedule construction mechanism. New sequences are generated by a variable neighborhood search method. Experimental and statistical analyses and results demonstrate that the proposed algorithm algorithm generates considerable cost savings when compared to the algorithms with only on-demand or reserved instances.
Long Chen 0021, Xiaoping Li 0001, Yucheng Guo, Rubén Ruiz
IEEE Trans. Cloud Comput.1
2020 A Metadata Inference Method for Building Automation Systems With Limited Semantic Information
abstract
Metadata in most existing building automation systems (BASs) is inconsistent, incomplete, and nondescriptive. This situation is a major obstacle to the widespread use of data analytics to improve the operation of buildings. In this article, we put forward a method to infer zone-level metadata from features derived from BAS data. The method includes two steps: 1) classification of BAS points into different types (e.g., indoor temperature, indoor temperature set point, airflow, airflow set point, damper position, and radiator valve position) and 2) association of BAS points based on their functional relationships (i.e., grouping the sensors, actuators, and set points of each zone together). The metadata inference method was demonstrated with data from zones served by four different air handling units (AHUs) in two office buildings in Ottawa, ON, Canada. The results from this case study indicate that common zone-level BAS point types can be accurately classified and associated even in the absence of intuitive data labels. Note to Practitioners-This article was motivated by the problem of metadata normalization in existing buildings, in order to scale up the application of smart building solutions in the real world. Existing metadata normalization approaches mainly focused on inferring the point types of the metadata with both semantic (label) and numerical information (time series readings). In this article, we put forward a method to infer zone-level metadata with numerical information only. Methods for both types of classification and relationships' association of the BAS points are investigated. The results from two office buildings indicate that the classification phase can achieve an average of 90% accuracy, while the association phase can obtain an average of 85% accuracy. The method was developed and demonstrated with a limited data set by using data exclusively from zone-level sensors, actuators, and set points. Future work is planned to extend the proposed method to more comprehensive BAS data sets with the system- and plant-level data as well.
Long Chen 0021, Burak Gunay, Zixiao Shi, Weiming Shen 0001, Xiaoping Li 0001
IEEE Trans Autom. Sci. Eng.1
2020 Resource Renting for Periodical Cloud Workflow Applications
abstract
Cloud computing is a new resource provisioning mechanism, which represents a convenient way for users to access different computing resources. Periodical workflow applications commonly exist in scientific and business analysis, among many other fields. One of the most challenging problems is to determine the right amount of resources for multiple periodical workflow applications. In this paper, the periodical workflow applications scheduling problem with total renting cost minimization is considered. The novelty of this work relies precisely on this objective function, which is more realistic in practice than the more commonly considered makespan minimization. An integer programming model is constructed for the problem under study. A Precedence Tree based Heuristic (PTH) is developed which considers three types of initial schedule construction methods. Based on the initial schedule, two improvement procedures are presented. The proposed methods are compared with existing algorithms for the related makespan based multiple workflow scheduling problem. Experimental and statistical results demonstrate the effectiveness and efficiency of the proposed algorithm.
Long Chen 0021, Xiaoping Li 0001, Rubén Ruiz
IEEE Trans. Serv. Comput.1
2018 Idle block based methods for cloud workflow scheduling with preemptive and non-preemptive tasks
Long Chen 0021, Xiaoping Li 0001, Rubén Ruiz
Future Gener. Comput. Syst.1
2018 Cloud workflow scheduling with hybrid resource provisioning
Long Chen 0021, Xiaoping Li 0001
J. Supercomput.1
2017 Cloud workflow scheduling with on-demand and spot block instances
abstract
Cloud computing enables users to access different resources conveniently based on the `pay-as-you-go' model. However, the unit cost of these on-demand instances are usually high. The spot instances provide a dynamic and cheaper manner for renting resources from the cloud. However, failures are often occurred due to the fluctuations of the price of the spot instance. It is a big challenge to determine the appropriate amounts of spot and on-demand resources in terms of users' requirements. In this paper, the workflow scheduling problem with both spot and on-demand instances is considered. The objective is to minimize the total renting cost under deadline constrains. An idle time block-based method is proposed to construct schedules for workflow applications. Schedules are improved by a forward and backward moving mechanism. Experimental and statistical results demonstrate the effectiveness of the proposed algorithm over a lot of tests with different sizes.
Long Chen 0021, Xiaoping Li 0001, Rubén Ruiz
CSCWD1
2016 Resources Renting with Reserved and On-Demand Instances for Cloud Workflow Applications
abstract
Cloud computing enables users to access different resources conveniently based on the "pay-as-you-go" model. However, the unit cost of this on-demand manner are usually higher than the reserved ones. Reallocating some high-usage on-demand instances to reserved instances can save considerable costs when renting resources from the cloud. It is a big challenge to determine the appropriate amount of reserved and on-demand instances in terms of users' requirements. In this paper, we consider deadline constrained workflow scheduling problems to minimize total renting costs with both reserved and on-demand instances. An integer programming model is constructed for the problem under study. A Precedence Tree based Heuristic (PTH) is developed which includes a dynamic initial schedule construction methods. Based on the initial schedule, an improvement procedure is presented. The proposed methods are compared with existing algorithms for the related makespan based workflow scheduling problem. Experimental and statistical results demonstrate the effectiveness and efficiency of the proposed algorithm.
Long Chen 0021, Xiaoping Li 0001
ICPADS1
2013 Bi-direction Adjust Heuristic for Workflow Scheduling in Clouds
abstract
This paper considers the workflow scheduling problem in Clouds with the hourly charging model and data transfer times. It deals with the allocation of tasks to suitable VM instances while maintaining the precedence constraints on one hand and meeting the workflow deadline on the other. A bi-direction adjust heuristic (BDA) is proposed for the considered problem. Matching of tasks and the VM types is modeled as Mixed Integer Linear programming (MILP) problem and solved using CPLEX at the first stage of BDA. In the second stage, forward and backward scheduling procedures are applied to allocate tasks to VM instances according to the result of the first stage. In the backward scheduling procedure, a priority rule considering the finish time, wasted time fractions and added hours is developed to make appropriate matches of tasks and free time slots. Extensive experimental results show that the proposed BDA heuristic outperforms the existing state-of-the-art heuristic ICPCP in all cases. Further, compared with ICPCP, about 80% percentage of VM renting cost is saved for instances with 900 tasks at most.
Zhicheng Cai, Xiaoping Li 0001, Long Chen 0021, Jatinder N. D. Gupta
ICPADS3
2012 Dynamic programming for services scheduling with start time constraints in distributed collaborative manufacturing systems
abstract
In this paper, the service scheduling problem with start time constraints is considered for distributed collaborative manufacturing systems, which is different from the discrete time-cost tradeoff problem (DTCTP), well studied during the past decades. The assumption that the ability of services is unlimited in DTCTP is seldom true for practical settings. The fact that most services have limited capabilities, especially for manufacturing services results in constraint start times for requirements. Such a DTCTP is modeled as the DTCTP-STC (discrete time-cost tradeoff problem with start time constraints), also proved to be NP-hard. A service is just available at some time point, which can be assigned as the start time negotiated between a broker and a provider. An effective dynamic programming algorithm is proposed with the time complexity O(N2Mv+1) for the DTCTP-STC. The impact of the number of nodes, the number of modes, and the complexity of the network on the computation time is analyzed by experiments. Simulated experiments are performs on randomly generated instances. The results illustrated that the proposal is very effective for small size instances. As well, the proposal is more suitable for the DTCTP-STC than the DTCTP with faster convergent speed for those general instances with fixed fewer modes.
Zhicheng Cai, Xiaoping Li 0001, Long Chen 0021
SMC3
2012 Heuristic methods for minimizing resource availability costs in multi-mode project scheduling
abstract
In this paper, a multi-mode project scheduling problem with deadline constraints is considered to minimize the resource availability cost. Modes of each activity are associated with different durations and renewable resources. Three kinds of rules are developed for activity selection, mode assignment, and time decision, respectively. A lot of combinations of the three kind rules are compared and the choosing probabilities are determined, based on which a regret probability based stochastic (RPBS) method is proposed for the considered problem. Computational results demonstrate that the RPBS method outperforms the existing one and the rule combinations in effectiveness but with a little more computation time.
Long Chen 0021, Xiaoping Li 0001, Zhicheng Cai
SMC1