Quan Zhou 0003

dblp:29/5849-3 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
12since 2021 · last 2025
0000-0001-6020-0416ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 8 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 KVPruner: Structural Pruning for Faster and Memory-Efficient Large Language Models
abstract
The bottleneck associated with the key-value(KV) cache presents a significant challenge during the inference processes of large language models. While depth pruning accelerates inference, it requires extensive recovery training, which can take up to two weeks. On the other hand, width pruning retains much of the performance but offers slight speed gains. To tackle these challenges, we propose KVPruner to improve model efficiency while maintaining performance. Our method uses global perplexity-based analysis to determine the importance ratio for each block and provides multiple strategies to prune non-essential KV channels within blocks. Compared to the original model, KVPruner reduces runtime memory usage by 50% and boosts throughput by over 35%. Additionally, our method requires only two hours of LoRA fine-tuning techniques on small datasets to recover most of the performance.
Quan Zhou 0003, Xuanang Ding, Yan Wang 0101, Zeming Ma
ICASSP2
2025 Priority-Aware Attention Meets Generative Flow Networks for Global Fixed-Priority Assignment
abstract
Ensuring timely task execution in multiprocessor systems under Global Fixed-Priority Scheduling (GFPS) remains a significant challenge. While traditional heuristic methods often struggle to derive feasible priority assignments for complex task sets, existing Deep Reinforcement Learning (DRL) approaches—though more powerful—suffer from two key limitations: (1) they fail to adequately capture relative priorities among high-priority tasks during encoding, and (2) their convergence to suboptimal policies due to sensitivity to local reward signals, an inherent drawback of reinforcement learning frameworks. To overcome these issues, we propose PAGFN, a novel learning-based framework that combines Priority-Aware Attention with Generative Flow Networks. We formulate the priority assignment problem as a Markov Decision Process (MDP) and employ an encoder-decoder architecture with auto-regressive decoding. Our framework introduces a priority-aware module to explicitly model task dependencies, enhancing assignment quality. Additionally, to mitigate sparse rewards, we integrate expert pretraining and prioritized experience replay, enabling diverse policy exploration without relying on intricate reward shaping. Experimental results on synthetic tasks show that PAGFN outperforms heuristic and learning-based baselines, particularly in scheduling previously unschedulable tasks, validating its effectiveness and scalability.
Shiwu Li, Jianjun Li 0010, Quan Zhou 0003
RTSS3
2025 Schedulability Analysis for Self-Suspending Tasks Under EDF-Like Scheduling
abstract
Real-time systems involve tasks that may voluntarily suspend their execution as they await specific events or resources. Such self-suspension can introduce further delays and unpredictability in scheduling, making the analysis more challenging. Most current schedulability analysis methods of self-suspending tasks focus on fixed-priority scheduling or tasks with constrained deadlines. This paper proposes two schedulability analysis methods for self-suspending tasks with arbitrary deadlines under earliest-deadline-first-like (EDF-like) scheduling. Both methods are designed for preemptive uniprocessor systems. We first present a jitter-based response time analysis (JRTA) method. JRTA is designed based on a self-suspending response time analysis (SS-RTA) method under earliest-deadline-first (EDF) scheduling. We first convert self-suspensions to release jitters and then present a response time analysis (RTA) method of tasks with release jitters under EDF-like scheduling. To address the complexity of JRTA, we propose an improved schedulability analysis (ISA), a sufficiency blocking-based method. Finally, we provide many simulation experiments under some EDF-like scheduling algorithms. The results verify the effectiveness and efficiency of both proposed methods.
Yan Wang 0101, Quan Zhou 0003, Junfei Li, Tan Tan 0004
IEEE Trans. Computers3
2025 Deadline and Period Assignment for Guaranteeing Timely Response of the Cyber-Physical System
abstract
Cyber-physical systems (CPSs) need to respond to each change of each monitored object in time. The entire response process can be divided into two stages: the update stage and the control stage. Tasks in CPSs can thus be divided into two kinds: update tasks and control tasks. Assigning deadlines and periods for tasks to ensure a timely response to each change of each monitored object is an important problem in CPS research. Existing methods can ensure that all changes of all objects can receive timely responses if tasks are schedulable. However, these methods have not made efforts to ensure the schedulability of tasks. Therefore, some tasks that can actually receive services cannot be serviced under these methods. In this article, we study the problem of assigning deadlines and periods for tasks while ensuring timely responses and maximizing the schedulability of tasks. Specifically, we find that the delayed deadline assignment for update tasks is the main factor that causes the low scheduling capability of existing methods. A new deadline and period assignment method is proposed based on an optimized deadline calculation scheme and an advanced deadline determination mechanism. Theoretical analysis proves the correctness and superiority of the proposed method. Experimental results show that the new method can improve 30.02% acceptance ratio and save 98.48% runtime on average, as compared to the state-of-the-art deadline and period assignment method.
Quan Zhou 0003, Si Cai, Jianjun Li 0010
ACM Trans. Design Autom. Electr. Syst.1
2022 Response Time Analysis for Hybrid Task Sets under Fixed Priority Scheduling
abstract
The task set in some real-time systems consists of adaptive variable-rate (AVR) tasks and time-triggered tasks. AVR tasks generate instances when the rotating source rotates to some special angles. Time-triggered tasks generate jobs irregularly under the constraint of minimum release time interval. Response time analysis (RTA) is an effective method to test the schedulability of tasks. The consideration on angular phase is conducive to improve the accuracy of RTAs. To the best of our knowledge, there is only one RTA method that considers angular phases and supports hybrid tasks. However, this method is incorrect since it may misjudge some unschedulable task sets as schedulable task sets. In this work, we propose the first correct RTA method that considers angular phases. The theoretical analysis proves the correctness of our method.
Quan Zhou 0003, Jihua Huang, Jianjun Li 0010
RTAS1
2022 SDNN: Symmetric deep neural networks with lateral connections for recommender systems
Runzhi Xu, Jianjun Li 0010, Guohui Li 0001, Peng Pan 0001, Quan Zhou 0003, Chaoyang Wang 0002
Inf. Sci.5
2022 FAS-DQN: Freshness-Aware Scheduling via Reinforcement Learning for Latency-Sensitive Applications
abstract
The demand for real-time data processing has become increasingly attractive in Cyber-Physical Systems(CPSs), especially for data-intensive embedded real-time applications. In order to timely perceive and respond to environmental changes, the basic design requirement in such systems is to provide data service with high freshness. As modern CPSs become more complex, there are a broad set of system mode switch behaviors, some unforeseen, in a dynamic computational environment. However, conventional control algorithms can hardly handle such new scenarios, since most of them assume that the operational behavior is fixed. In this paper, we study the problem of how to maximize the freshness of data in multi-modal systems. We first use a recently proposed new conception, namely Age of Information (AoI) to quantify the freshness of data by combining the AoI metric with real-time constraints. Then, we propose, to our knowledge, the first freshness-aware scheduling solution to settle the problem via deep reinforcement learning(RL). To be specific, we develop an RL framework that can continuously update its scheduling strategies and maximize the freshness of data in the long term. Extensive simulation experiments are conducted and the results demonstrate that the proposed FAS-DQN outperforms other traditional state-of-the-art methods in terms of data freshness.
Chunyang Zhou, Guohui Li 0001, Jianjun Li 0010, Quan Zhou 0003, Bing Guo 0003
IEEE Trans. Computers4
2022 An Efficient Execution Framework of Two-Part Execution Scenario Analysis
abstract
Response Time Analysis ( RTA ) is an important and promising technique for analyzing the schedulability of real-time tasks under both Global Fixed-Priority ( G-FP ) scheduling and Global Earliest Deadline First ( G-EDF ) scheduling. Most existing RTA methods for tasks under global scheduling are dominated by partitioned scheduling, due to the pessimism of the -based interference calculation where is the number of processors. Two-part execution scenario is an effective technique that addresses this pessimism at the cost of efficiency. The major idea of two-part execution scenario is to calculate a more accurate upper bound of the interference by dividing the execution of the target job into two parts and calculating the interference on the target job in each part. This article proposes a novel RTA execution framework that improves two-part execution scenario by reducing some unnecessary calculation, without sacrificing accuracy of the schedulability test. The key observation is that, after the division of the execution of the target job, two-part execution scenario enumerates all possible execution time of the target job in the first part for calculating the final Worst-Case Response Time ( WCRT ). However, only some special execution time can cause the final result. A set of experiments is conducted to test the performance of the proposed execution framework and the result shows that the proposed execution framework can improve the efficiency of two-part execution scenario analysis by up to in terms of the execution time.
Ding Han, Guohui Li 0001, Quan Zhou 0003, Jianjun Li 0010, Xiaofei Hu
ACM Trans. Design Autom. Electr. Syst.3
2022 Double Attention Convolutional Neural Network for Sequential Recommendation
abstract
The explosive growth of e-commerce and online service has led to the development of recommender system. Aiming to provide a list of items to meet a user’s personalized need by analyzing his/her interaction 1 history, recommender system has been widely studied in academic and industrial communities. Different from conventional recommender systems, sequential recommender systems attempt to capture the pattern of users’ sequential behaviors and the evolution of users’ preferences. Most of the existing sequential recommendation models only focus on user interaction sequence, but neglect item interaction sequence. An item interaction sequence also contains rich contextual information for capturing the item’s dynamic characteristic, since an item’s dynamic characteristic can be reflected by the users who interact with it in a period. Furthermore, existing dual sequential models use the same method to handle the user interaction sequence and item interaction sequence, and do not consider their different characteristics. Hence, we propose a novel D ouble A ttention C onvolution N eural N etwork (DACNN) , which incorporates user interaction sequence and item interaction sequence into an integrated neural network framework. DACNN leverages the strength of attention mechanism to capture the temporary suitability and adopts CNN to extract local sequential features. Experimental evaluations on the real datasets show that DACNN outperforms the baseline approaches.
Qi Chen 0017, Guohui Li 0001, Quan Zhou 0003, Deqing Zou
ACM Trans. Web3
2021 Limited Busy Periods in Response Time Analysis for Tasks Under Global EDF Scheduling
abstract
Response time analysis (RTA) is a kind of effective methods to test the schedulability of real-time tasks under the global earliest deadline first (G-EDF) scheduling. The main idea of the RTA methods is to judge the schedulability of the task set by comparing the worst-case response time (WCRT) and deadline of each task. For obtaining a safe WCRT, existing RTA methods assume that the response times of carry-in jobs are equal to the WCRTs of the tasks generating these jobs. This assumption causes the low accuracy of existing RTA methods, since the busy period lengths of carry-in jobs are usually limited and the response times of these jobs thus cannot always reach the assumed WCRTs. This article proposes a new RTA method that has a higher accuracy than existing RTAs. The main idea of the new method is to limit the busy period lengths of carry-in jobs in the WCRT calculation. We also propose an efficiency improvement scheme for the new method. A set of experiments is conducted to evaluate the performances of the new method and the efficiency improvement scheme. The experimental results show that the new method has a higher acceptance ratio than the state-of-the-art RTA methods and the efficiency improvement scheme can significantly reduce the execution time of the new method.
Quan Zhou 0003, Guohui Li 0001, Chunyang Zhou, Jianjun Li 0010
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2021 Guaranteeing Timely Response to Changes of Monitored Objects by Assigning Deadlines and Periods to Tasks
abstract
Timely response to changes of monitored objects is the key to ensuring the safety and reliability of cyber-physical systems (CPSs). There are two kinds of tasks in CPSs: update tasks and control tasks. Update tasks are responsible for updating the data in the system based on the state of the objects they monitor. Control tasks are responsible for making decisions based on the data in the system. The response time of the system to the change of a monitored object consists of two parts: the time taken by update tasks to reflect the change to the system, and the time taken by control tasks to make decisions according to the data in the system. Deadlines and periods of update tasks and control tasks directly affect the response time. Reasonable deadline and period assignment is the key to ensuring timely response to the changes of monitored objects. In this paper, we study the deadline and period assignment in CPSs. To the best of our knowledge, all existing work only focuses on the deadline and period assignment for update tasks with the goal of ensuring the freshness of the data in CPSs, and this is the first study focusing on the deadline and period assignment for both update tasks and control tasks with the goal of ensuring timely response to the changes of monitored objects. A new problem about response time control and system workload control is defined in this paper. Two deadline and period assignment methods are proposed to solve the defined problem. All the proposed methods can be used in the CPSs adopting the earliest deadline first (EDF) scheduling method. Experiments with randomly generated tasks are conducted to evaluate the performance of the proposed methods in terms of acceptance ratio and execution efficiency.
Quan Zhou 0003, Guohui Li 0001, Qi Chen 0017, Jianjun Li 0010
ACM Trans. Embed. Comput. Syst.1
2021 Excluding Parallel Execution to Improve Global Fixed Priority Response Time Analysis
abstract
Response Time Analysis (RTA) is an effective method for testing the schedulability of real-time tasks on multiprocessor platforms. Existing RTAs for global fixed priority scheduling calculate the upper bound of the worst case response time of each task. Given a target task, existing RTAs first calculate the workload upper bound of each higher priority task (than the target task), and then calculate the interference on the target task by each higher priority task according to the obtained workload upper bounds. The workload of a task consists of three parts: carry-in, body and carry-out. The interference from all these three parts may be overestimated in existing RTAs. However, although the overestimation of the interference from body is the major factor that causes the low accuracy of existing RTAs, all existing work only focuses on how to reduce the overestimation of the interference from carry-in, and there is no method to reduce the overestimation of the interference from body or carry-out. In this work, we propose a method to calculate the lower bound of the accumulative time in which the target task and higher priority tasks are executed in parallel. By excluding the parallel execution time from the interference, we derive a new RTA test that can reduce the overestimation of the interference from all three parts of the workload. Extensive experiments are conducted to verify the superior performance of the proposed RTA test.
Quan Zhou 0003, Jianjun Li 0010, Guohui Li 0001
ACM Trans. Embed. Comput. Syst.1
2020 DDFL: A Deep Dual Function Learning-Based Model for Recommender Systems
Syed Tauhid Ullah Shah, Jianjun Li 0010, Zhiqiang Guo, Guohui Li 0001, Quan Zhou 0003
DASFAA (3)5
2019 Response Time Analysis for Tasks with Fixed Preemption Points under Global Scheduling
abstract
As an effective method for detecting the schedulability of real-time tasks on multiprocessor platforms, Response time analysis (RTA) has been deeply researched in recent decades. Most of the existing RTA methods are designed for tasks that can be preempted at any time. However, in some real-time systems, a task may have some fixed preemption points (FPPs) that divide its execution into a series of non-preemptive regions (NPRs). In such environments, the task can only be preempted at its FPPs, which makes existing RTA methods for arbitrary preemption tasks not applicable. In this article, we study the schedulability analysis on tasks with FPPs under both global fixed-priority (G-FP) scheduling and global earliest deadline first (G-EDF) scheduling. First, based on the idea of limiting the time interval between two consecutive executions of an NPR, a novel RTA method for tasks with FPPs under G-FP scheduling is proposed. Second, we propose an effective RTA method for tasks with FPPs under G-EDF scheduling. Finally, extensive simulations are conducted and the results validate the effectiveness of the proposed methods.
Quan Zhou 0003, Guohui Li 0001, Jianjun Li 0010, Chenggang Deng, Ling Yuan
ACM Trans. Embed. Comput. Syst.1
2018 Execution-Efficient Response Time Analysis on Global Multiprocessor Platforms
abstract
Response time analysis (RTA) is an important and fundamental tool for analyzing the schedulability of real-time tasks on multiprocessor platforms, and many promising techniques have been developed during the past few years. However, most of the existing researches focus on improving the analysis precision, while less has been done on enhancing the execution efficiency. In this paper, we take, to our best knowledge, the first effort towards improving the efficiency of the state-of-the-art RTA methods for sporadic tasks under both global fixed-priority (G-FP) and global earliest deadline first (G-EDF) scheduling. Specifically, three factors that impact the efficiency of existing global RTA tests are identified: 1) pessimistic initial value when computing the worst-case response time (WCRT) of tasks under both G-FP and G-EDF; 2) conservative interference upper bound under both G-FP and G-EDF; and 3) unnecessary recalculation of WCRTs when there is update of any task WCRT under G-EDF. By addressing these three limitations, we propose two efficient RTA methods for G-FP and G-EDF scheduling, respectively, which achieve better run-time performance but without sacrificing any analysis precision. Experimental evaluations with randomly generated task sets show that the proposed methods exhibit remarkable performance improvements and can save on average 60 and 61 percent run time, as compared to the state-of-the-art technologies under G-FP and G-EDF scheduling, respectively.
Quan Zhou 0003, Guohui Li 0001, Jianjun Li 0010, Chenggang Deng
IEEE Trans. Parallel Distributed Syst.1
2017 Dynamic priority scheduling of periodic queries in on-demand data dissemination systems
Quan Zhou 0003, Guohui Li 0001, Jianjun Li 0010, LihChyun Shu, Cong Zhang 0007, Fumin Yang
Inf. Syst.1
2017 Deadline and Period Assignment for Update Transactions in Co-Scheduling Environment
abstract
Deriving deadline and period for update transactions to maintain temporal consistency has long been recognized as an important problem in real-time database research. Despite years of active study, most of the past work only focuses on the scheduling of update transactions, and neglects the impact of control transactions by assuming anUpdate Firstpolicy where control transactions are always assigned lower priorities than the update transactions. On the other hand, most existing work on co-scheduling of update and control transactions has been focused on meeting the deadlines of all the control transactions while maximizing the quality of data of the real-time data objects. In this paper, we study the co-scheduling problem of update and control transactions by satisfying the deadline constraints of control transactions and the temporal validity constraints of update transactions simultaneously. Specifically, we consider the problem of how to derive deadline and period for update transactions to maintain the temporal consistency of real-time data objects, while guaranteeing the hybrid transaction set to be EDF-schedulable. To address this problem, we first borrow the idea from$\mathsf{minD}$[14]to derive a solution called$\mathsf{minD}^\ast$, which can compute deadline and period for update transactions effectively. Next, based on a sufficient condition to derive the minimum possible deadline for each update transaction, we propose a more efficient algorithmMinimum Deadline Calculation($\mathsf{MDC}$), which can guarantee to derive a solution, given that one does exist. Finally, the effectiveness and efficiency of the proposed algorithms are validated through extensive simulation experiments.
Guohui Li 0001, Chenggang Deng, Jianjun Li 0010, Quan Zhou 0003, Wei Wei 0002
IEEE Trans. Computers4
2017 Improved Carry-in Workload Estimation for Global Multiprocessor Scheduling
abstract
As an important and fundamental tool for analyzing the schedulability of a real-time task set on the multiprocessor platform, response time analysis (RTA) has been researched for several years on both Global Fixed Priority (G-FP) and Global Earliest Deadline First (G-EDF) scheduling. This paper proposes a new analysis that improves over current state-of-the-art RTA methods for both G-FP and G-EDF scheduling, by reducing their pessimism. The key observation is that when estimating the carry-in workload, all the existing RTA techniques depend on the worst case scenario in which the carry-in job should execute as late as possible and just finishes execution before its worst case response time (WCRT). But the carry-in workload calculated under this assumption may be over-estimated, and thus the accuracy of the response time analysis may be impacted. To address this problem, we first propose a new method to estimate the carry-in workload more precisely. The proposed method does not depend on any specific scheduling algorithm and can be used for both G-FP and G-EDF scheduling. We then propose a general RTA algorithm that can improve most existing RTA tests by incorporating our carry-in estimation method. To further improve the execution efficiency, we also introduce an optimization technique for our RTA tests. Experiments with randomly generated task sets are conducted and the results show that, compared with the state-of-the-art technologies, the proposed tests exhibit considerable performance improvements, up to 9 and 7.8 percent under G-FP and G-EDF scheduling respectively, in terms of schedulability test precision.
Quan Zhou 0003, Guohui Li 0001, Jianjun Li 0010
IEEE Trans. Parallel Distributed Syst.1
2015 A Novel Scheduling Algorithm for Supporting Periodic Queries in Broadcast Environments
abstract
Being a proven efficient approach to answering queries that have common data needs, data broadcast has received much attention in the past decade, especially for dynamic and large-scale data dissemination. An important class of emerging data broadcast applications must monitor multiple data items continuously in order to enable data-driven decision making. For such applications, an important problem that must be addressed is how to disseminate data to periodic continuous queries so that all the requests can be satisfied while the bandwidth utilization is minimized. To our best knowledge, the only known work on this topic is the RM-UO algorithm proposed in the work of Huang et al. (2012). However, the RM-UO algorithm simply utilizes the Sr algorithm introduced in the work of Han et al. (1996) to transform the original queries into 2-harmonic tasks, which would lead to a considerable waste of available bandwidth. In this paper, based on the observation that some queries can be merged to save bandwidth consumption, we propose two merging polices namely Multiple Query Merging (MQM) and Redundant Query Merging (RQM), and show that both can lead to notable bandwidth savings. Further, to disseminate data to periodic continuous queries, we implement a unified scheduling algorithm called UM, which combines both MQM and RQM. Extensive experiments have been conducted to compare our UM algorithm with RM-UO, and the results show that UM outperforms RM-UO considerably in terms of wireless bandwidth consumption and query service ratio.
Guohui Li 0001, Quan Zhou 0003, Jianjun Li 0010
IEEE Trans. Mob. Comput.2