VLDB 2026 Research / reviewers in the wild / expert
Changyao Lin
dblp:305/9647
· DBLP profile ↗
14ranked-venue papers
10as first author
14since 2021 · last 2026
0000-0001-6805-2649ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 8 first-author · 10 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Coconut: Multilevel Collaborative Deployment for Real-Time Deep Learning Tasks in Heterogeneous Edge GPU Cluster
Changyao Lin, Zhenming Chen, Jie Liu 0001 |
IEEE Internet Things J. | 1 |
| 2026 | Cola: Cross-Processor Operator Parallelism for Asynchronous Deep Learning Inference
Changyao Lin, Zhenming Chen, Jie Liu 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | E4: Energy-Efficient DNN Inference for Edge Video Analytics via Early Exiting and DVFSabstractDeep neural network (DNN) models are increasingly popular in edge video analytic applications. However, the computeintensive nature of DNN models pose challenges for energyefficient inference on resource-constrained edge devices. Most existing solutions focus on optimizing DNN inference latency and accuracy, often overlooking energy efficiency. They also fail to account for the varying complexity of video frames, leading to sub-optimal performance in edge video analytics. In this paper, we propose an EnergyEfficient Early-Exit (E4) framework that enhances DNN inference efficiency for edge video analytics by integrating a novel early-exit mechanism with dynamic voltage and frequency scaling (DVFS) governors. It employs an attentionbased cascade module to analyze video frame diversity and automatically determine optimal DNN exit points. Additionally, E4 features a just-in-time (JIT) profiler that uses coordinate descent search to co-optimize CPU and GPU clock frequencies for each layer before the DNN exit points. Extensive evaluations demonstrate that E4 outperforms current state-of-the-art methods, achieving up to 2.8× speedup and 26% average energy saving while maintaining high accuracy. Yang Zhao 0020, Ming-Ching Chang, Changyao Lin, Jie Liu 0001 |
AAAI | 4 |
| 2025 | E3: Early Exiting with Explainable AI for Real-Time and Accurate DNN Inference in Edge-Cloud SystemsabstractEdge intelligence applications frequently generate deep learning inference tasks with varying Service Level Objectives (SLO, such as accuracy and real-time requirements). For such tasks, recent progressive inference modes support early exit from different branches to satisfy inference requirements. However, existing edge-cloud progressive neural architectures cannot simultaneously achieve high accuracy and real-time performance for different data features. Therefore, we utilize explainable AI technique to construct and train a novel progressive neural architecture E3. E3 can progressively extract the most important features for inference, ensuring higher accuracy at early-exit points. While the less important features in the later stage are highly compressible, thereby reducing edge-cloud transmission overhead. Furthermore, E3 cooperates with online execution control to launch tasks and decide the exit point for each task, ensuring resource utilization and real-time performance, and adapting to bandwidths and deadlines. Experimental results on various edge-cloud platforms, datasets, and reference models demonstrate that E3 is more lightweight, efficient, energy-saving, and incurs almost no additional runtime overhead compared to traditional architectures. Under stringent deadlines, the average accuracy of tasks increases by > 50%, and the deadline satisfaction rate approaches 100%. Changyao Lin, Zhenming Chen, Jie Liu 0001 |
SenSys | 1 |
| 2025 | Exploiting Operator-Level Concurrency Control to Guide Deployment for Real-Time Tasks in Edge AI ClusterabstractExisting task deployment frameworks for edge clusters optimize at the model-level, lacking fine-grained resource awareness and concurrency control, where the urgent tasks are frequently blocked and miss their deadlines. Therefore, we propose a multi-level collaborative deployment framework Coconut for real-time deep learning tasks in the typical heterogeneous edge GPU cluster. Coconut collaboratively optimizes model deployment and fine-grained concurrency control. To address the high complexity of multi-level collaborative optimization, we employ an efficient learning-based search algorithm. Based on the operator-level information, we also pre-train an accurate latency predictor for each device, enabling centralized optimization to further accelerate the search. We conduct a preliminary evaluation in an edge cluster and validate the effectiveness of Coconut. Changyao Lin, Jie Liu 0001 |
SenSys | 1 |
| 2025 | TOP: Task-Based Operator Parallelism for Asynchronous Deep Learning Inference on GPUabstractCurrent deep learning compilers have made significant strides in optimizing computation graphs for single- and multi-model scenarios. However, they lack specific optimizations for asynchronous multi-task inference systems. In such systems, tasks arrive dynamically, leading to diverse inference progress for each model. This renders traditional optimization strategies based solely on the original computation graph suboptimal or even invalid. Furthermore, existing operator scheduling methods do not account for parallel task pipelines involving the same model. Task pipelines present additional opportunities for optimization. Therefore, we propose Task-based Operator Parallelism (TOP). TOP incorporates an understanding of the impact of task arrival patterns on the inference progress of each model. It leverages the multi-agent reinforcement learning algorithm MADDPG to cooperatively optimize the task launcher and model scheduler, generating an optimal pair of dequeue frequency and computation graph. The objective of TOP is to enhance resource utilization, increase throughput, and allocate resources judiciously to prevent task backlog. To expedite the optimization process in TOP, we introduce a novel stage partition method using the GNN-based Policy Gradient (GPG) algorithm. Through extensive experiments on various devices, we demonstrate the efficacy of TOP. It outperforms the state-of-the-art in operator scheduling for both single- and multi-model task processing scenarios. Benefiting from TOP, we can significantly enhance the throughput of a single model by increasing its concurrency or batch size, thereby achieving self-acceleration. Changyao Lin, Zhenming Chen, Jie Liu 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2024 | StressViT: Splitting and Compressing Vision Transformer Through Edge-Cloud Collaboration
Changyao Lin, Yi Liu 0096, Chengxiang Li, Hao Zhang 0016, Jing Jin 0003, Jie Liu 0001 |
ICPR (5) | 1 |
| 2024 | Poster Abstract: Xpi: Real-Time Progressive Inference Serving with Explainable AI in Edge-Cloud SystemsabstractThe constrained computing and memory resources at the edge pose challenges for satisfying different service-level objectives (SLOs) of deep learning inference requests. In this paper, we propose a novel edge-cloud progressive inference framework Xpi, which integrates explainable AI technique to facilitate early-exit, and learning-based online execution control to satisfy different SLOs and optimize edge resource overheads. We implement Xpi on an edge-cloud platform, and conduct partial experiments on two datasets. Xpi outperforms several advanced edge-cloud progressive inference frameworks in terms of accuracy and deadline satisfaction rate. Changyao Lin, Zhenming Chen, Jie Liu 0001 |
IPSN | 1 |
| 2024 | COS: Cross-Processor Operator Scheduling for Multi-Tenant Deep Learning InferenceabstractMulti-tenant inference, as a prevalent inference paradigm nowadays, requires deploying multiple deep learning models on the hardware platform to concurrently process inference tasks. Modern platforms are typically equipped with various heterogeneous processors, such as CPU-GPU platform. To reduce resource contention and improve Quality of Service (QoS) in the multi-tenant scenario, existing work has studied cross-processor inference at the model- and layer-level. However, coarse-grained scheduling cannot flexibly account for subtle resource fluctuations, which may lead to task blockages and incur significant processor switching overheads. Such work usually requires extensive modification and retraining of the models. Therefore, we propose a finer-grained operator-level cross-processor scheduling framework COS, which can more precisely divide the computational workloads and switching overheads for the tenants, without modifying or retraining. We introduce a novel intermediate representation to abstract and simplify the scheduling problem, and propose an efficient two-phase search algorithm. COS is automated and easy-to-scale, through experiments on various heterogeneous hardware platforms and models, we demonstrate that COS is more flexible and effective than layer-level scheduling, and achieves higher throughput than single-processor processing in the multi-tenant scenario. Furthermore, COS is an offline optimization method, and its overhead is highly acceptable. Changyao Lin, Jie Liu 0001 |
IWQoS | 1 |
| 2024 | A QoS-Aware Training Framework for ViT Compression, Partition, and DistillationabstractIn this paper, we jointly optimize compression, partition and distillation for visual Transformer. We analyze the relationship among the three modules and integrate them into a QoS-aware training framework. By coordinating the model compression, edge-cloud partition, and knowledge distillation during training, the model architecture and accuracy are optimized simultaneously. The framework considers the differences in computing power and memory between edge and cloud, and can trade off QoS metrics such as the memory overhead, end-to-end latency, accuracy at multiple granularities. Changyao Lin, Chengxiang Li, Jie Liu 0001 |
IWQoS | 1 |
| 2024 | DVFO: Learning-Based DVFS for Energy-Efficient Edge-Cloud Collaborative InferenceabstractDue to limited resources on edge and different characteristics of deep neural network (DNN) models, it is a big challenge to optimize DNN inference performance in terms of energy consumption and end-to-end latency. In addition to dynamic voltage frequency scaling (DVFS) technique, edge-cloud architecture provides a collaborative approach for efficient DNN inference. However, current edge-cloud collaborative inference methods have not optimized various compute resources on edge devices. Thus, we propose DVFO, a novel DVFS-enabled edge-cloud collaborative inference framework, which co-optimizes DVFS and offloading parameters via deep reinforcement learning (DRL). Specifically, DVFO automatically co-optimizes 1) the CPU, GPU and memory frequencies of edge devices, and 2) the offloaded feature map. In addition, it leverages athinking-while-movingconcurrent mechanism to accelerate the DRL learning process, and aspatial-channel attentionmechanism to identify the less important DNN feature map for efficient offloading. This approach improves inference performance for different DNN models under various edge-cloud network conditions. Extensive evaluations using two datasets and six widely-deployed DNN models on five heterogeneous edge devices show that DVFO significantly reduces the energy consumption by 33% on average, compared to state-of-the-art schemes. Moreover, DVFO achieves up to 28.6%∼59.1% end-to-end latency reduction, while maintaining accuracy within 1% loss on average. Yang Zhao 0020, Changyao Lin, Jie Liu 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2023 | POS: An Operator Scheduling Framework for Multi-model Inference on Edge Intelligent ComputingabstractEdge intelligent applications, such as autonomous driving usually deploy multiple inference models on resource-constrained edge devices to execute a diverse range of concurrent tasks, given large amounts of input data. One challenge is that these tasks need to produce reliable inference results simultaneously with millisecond-level latency to achieve real-time performance and high quality of service (QoS). However, most of the existing deep learning frameworks only focus on optimizing a single inference model on an edge device. To accelerate multi-model inference on a resource-constrained edge device, in this paper we propose POS, a novel operator-level scheduling framework that combines four operator scheduling strategies. The key to POS is a maximum entropy reinforcement learning-based operator scheduling algorithm MEOS, which generates an optimal schedule automatically. Extensive experiments show that POS outperforms five state-of-the-art inference frameworks: TensorFlow, PyTorch, TensorRT, TVM, and IOS, by up to 1.2 × ∼ 3.9 × inference speedup consistently, with 40% improvement on GPU utilization. Meanwhile, MEOS reduces the scheduling overhead by 37% on average, compared to five baseline methods including sequential execution, dynamic programming, greedy scheduling, actor-critic, and coordinate descent search algorithms. Yang Zhao 0020, Changyao Lin, Jie Liu 0001 |
IPSN | 4 |
| 2021 | Choosing Appropriate AI-enabled Edge Devices, Not the Costly OnesabstractAdvances in Edge AI make it possible to achieve inference deep learning for emerging applications, e.g., smart transportation and smart city on the edge in real-time. Nowadays, different industry companies have developed several edge AI devices with various architectures. However, it is hard for application users to justify how to choose the appropriate edge-AI, due to the lack of benchmark testing results and testbeds specifically used to evaluate the system performance for those edge-AI systems. In this paper, we attempt to design a benchmark test platform for the edge-AI devices and evaluate six mainstream edge devices that are equipped with different computing powers and AI chip architectures. Throughput, power consumption ratio, and cost-effectiveness are chosen as the performance metrics for the evaluation process. Three classic deep learning workloads: object detection, image classification, and natural language processing are adopted with different batch sizes. The results show that under different batch sizes, compared with traditional edge devices, edge devices equipped with AI chips have out-performance in throughput, power consumption ratio, and cost-effectiveness by 134×, 57×, and 32×, respectively. From system perspective, our work not only demonstrates the effective AI capabilities of those edge AI devices, but also provide suggestions for AI optimization at edge in details. Changyao Lin, Shihui Wen, Jie Liu 0001 |
ICPADS | 3 |
| 2021 | ECSRL: A Learning-Based Scheduling Framework for AI Workloads in Heterogeneous Edge-Cloud SystemsabstractRecent advances in both lightweight models and edge computing make it possible for inference tasks to be executed concurrently on resource-constrained edge devices. However, our preliminary experiments show that the execution of different lightweight models on edge devices may lead to a performance downgrade. In this paper, we propose a Learning-Based Scheduling Framework---ECSRL, to optimize the latency and power consumption for those inference tasks running in heterogeneous Edge-Cloud systems. Changyao Lin, Jie Liu 0001 |
SenSys | 1 |