VLDB 2026 Research / reviewers in the wild / expert
Yaqiong Peng
dblp:120/1801
· DBLP profile ↗
15ranked-venue papers
9as first author
3since 2021 · last 2025
0000-0002-8482-3631ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 7 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Decentralized Request Dispatch for Edge-Clouds: A Diffusion-Based Reinforcement Learning ParadigmabstractEdge-cloud systems have the potential to achieve ubiquitous computing by providing services in close proximity to users that submit service requests. The key challenge is how to efficiently orchestrate services and dispatch requests to satisfy the Quality of Service (QoS) requirements of users in dynamic edge-cloud environments. With the benefit of efficiently adapting to uncertainty, Reinforcement Learning (RL) based approaches are proposed to solve the request dispatch problem in edge-cloud environments. However, existing RL based approaches are often constrained by inexpressive policies that make highly suboptimal decisions in the field of request dispatch for edge-clouds. To enhance the effectiveness of RL in guaranteeing the QoS requirements of users, this paper presents D2Sched, a novel scheduling framework that represents the policy networks of Multi-Agent Deep Reinforcement Learning (MADRL) as diffusion models to generate request dispatch decisions. To improve the valid probability of generated dispatch decisions, D2Sched coordinates all agents for the resource competition among different requests by carefully considering the availability of system resources and latency targets of requests. Extensive experiments using synthetic and real traces demonstrate that D2Sched can improve the average system throughput under QoS requirements of users by up to 20.1% compared to representative baselines. Yaqiong Peng, Haocheng Peng |
IEEE Trans. Serv. Comput. | 1 |
| 2024 | InferFair: Towards QoS-aware scheduling for performance isolation guarantee in heterogeneous model serving systems
Yaqiong Peng, Haocheng Peng |
Future Gener. Comput. Syst. | 1 |
| 2024 | Serving DNN Inference With Fine-Grained Spatio-Temporal Sharing of GPU ServersabstractDeep Neural Networks(DNNs) are commonly deployed as online inference services. To meet interactive latency requirements of requests, DNN services require the use ofGraphics Processing Unit(GPU) to improve their responsiveness. The unique characteristics of inference workloads pose new challenges to manage GPU resources. First, the GPU scheduler needs to carefully manage requests to meet their latency targets. Second, a single inference task often underutilizes GPU resources. Third, the fluctuating patterns of inference workloads pose difficulties in determining the resources allocated to each DNN model. Therefore, it is critical for the GPU scheduler to maximize GPU utilization by collocating multiple DNN models without violating the latencyService-Level Objectives(SLOs) of requests. However, we find that existing works are not adequate for achieving this goal among latency-sensitive inference tasks. Hence, we propose FineST, a scheduling framework for serving DNNs with fine-grained spatio-temporal sharing of GPU inference servers. To maximize GPU utilization, FineST allocates intra-GPU computing resources from both spatial and temporal dimensions across DNNs in a cost-effective way, while predicting interference overheads under diverse consolidated executions for controlling SLO violation rates. Compared to a state-of-the-art work, FineST improves the peak throughput of serving heterogeneous DNNsby up to 64.7% under SLO constraints. Yaqiong Peng, Weiguo Gao, Haocheng Peng |
IEEE Trans. Serv. Comput. | 1 |
| 2020 | CHEAPS2AGA: Bounding Space Usage in Variance-Reduced Stochastic Gradient Descent over Streaming Data and Its Asynchronous Parallel Variants
Yaqiong Peng, Haiqiang Fei, Zhenquan Ding, Zhiyu Hao |
ICA3PP (2) | 1 |
| 2020 | pRnR: A Parallel Record-Replay Framework for Virtual MachinesabstractThe record and replay(RnR) technology of virtual machine(VM) provides the ability to reproduce the past execution of a VM deterministically. It has many promising applications in the cloud environment, including fault tolerance, security analysis, and failure diagnosis. Existing studies in this area pay more effort in optimizing the record method, such as reducing performance penalty and storage costs. However, considering that many practical applications follow the record once, replay many mode, the optimization for the replay is more critical, especially for efficiency. In this paper, we propose pRnR, a novel parallel RnR framework, to support efficient replay. By combining the native RnR framework with an improved continuous snapshots mechanism, pRnR divides the full execution into many independent and complete slices, each of which supports arbitrary replay. In addition, it supports two replay modes to improve replay efficiency, i.e., multi-slice parallel replay and multi-dimension parallel replay. Moreover, we apply our pRnR framework to syscall-based diagnosis to demonstrate its usability. The experimental results show that pRnR is more efficient than existing RnR frameworks. Wei Wang 0428, Lei Cui 0003, Zhiyu Hao, Haiqiang Fei, Chonghua Wang, Yaqiong Peng |
ICCD | 6 |
| 2020 | Lock-Free Parallelization for Variance-Reduced Stochastic Gradient Descent on Streaming DataabstractStochastic Gradient Descent (SGD) is an iterative algorithm for fitting a model to the training dataset in machine learning problems. With low computation cost, SGD is especially suited for learning from large datasets. However, the variance of SGD tends to be high because it uses only a single data point to determine the update direction at each iteration of gradient descent, rather than all available training data points. Recent research has proposed variance-reduced variants of SGD by incorporating a correction term to approximate full-data gradients. However, it is difficult to parallelize such variants with high performance and accuracy, especially on streaming data. As parallelization is a crucial requirement for large-scale applications, this article focuses on the parallel setting in a multicore machine and presents LFS-STRSAGA, a lock-free approach to parallelizing variance-reduced SGD on streaming data. LFS-STRSAGA embraces a lock-free data structure to process the arrival of streaming data in parallel, and asynchronously maintains the essential information to approximate full-data gradients with low cost. Both our theoretical and empirical results show that LFS-STRSAGA matches the accuracy of the state-of-the-art variance-reduced SGD on streaming data under sparsity assumption (common in machine learning problems), and that LFS-STRSAGA reduces the model update time by over 98 percent. Yaqiong Peng, Zhiyu Hao, Xiao-chun Yun |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2019 | Fast Wait-Free Construction for Pool-Like Objects with Weakened Internal Order: Stacks as an ExampleabstractThis paper focuses on a large class of concurrent data structures that we call pool-like objects (e.g., stack, double-ended queue, and queue). Performance and progress guarantee are two important characteristics for concurrent data structures. In the aspect of performance, weakening the internal order in a pool-like object is an effective technique to reduce the synchronization cost among threads accessing the object, but no objects with weakened internal order provide a progress guarantee as strong as wait-freedom. Meanwhile, wait-free algorithms tend to be inefficient, which is mainly attributed to the helping mechanisms. Based on the philosophy of existing helping mechanisms, a wait-free pool-like object with weakened internal order would suffer from unnecessary process of getting the latest object state and synchronization. This paper takes a state-of-the-art implementation of stacks with weakened internal order as an example, and transforms it into a highly-efficient wait-free stack named WF-TS-Stack. The transformation method includes a helping mechanism with state reuse and a relaxed removal scheme. In addition, we use a simple and effective scheme to further improve the performance of WF-TS-Stack in Non-Uniform Memory Access (NUMA) architectures. Our evaluation with representative benchmarks shows that WF-TS-Stack outperforms its original building blocks by up to 1.45× at maximum concurrency. We also discuss how to yield an efficient double-ended queue (deque) variant of WF-TS-Stack, because deque is a more generalized pool-like object. Yaqiong Peng, Xiao-chun Yun, Zhiyu Hao |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | FA-Stack: A Fast Array-Based Stack with Wait-Free Progress GuaranteeabstractThe prevalence of multicore processors necessitates the design of efficient concurrent data structures. Shared concurrent stacks are widely used as inter-thread communication structures in parallel applications. Wait-free stacks can ensure that each thread completes operations on them in a finite number of steps. This characteristic is valuable for parallel applications and operating systems, especially in real-time environments. Unfortunately, because wait-free algorithms are typically hard to design and considered inefficient, practical wait-free stacks are rare. In this paper, we present a practical, fast array-based concurrent stack with wait-free progress guarantee, named FA-Stack. A series of optimizations are proposed to bound the number of steps required to complete every push and pop operation. In addition, FA-Stack adopts a time-stamped scheme to reclaim memory. We use linearizability, a correctness condition for concurrent data structures, to prove that FA-Stack is a wait-free linearizable stack with respect to the Last in First Out (LIFO) semantics. Our evaluation with representative benchmarks shows that FA-Stack is an efficient wait-free stack. For example, compared to Sim-Stack (a state-of-the-art wait-free stack), FA-Stack improves the throughput of halfhalf benchmark by upto 2.4×. Yaqiong Peng, Zhiyu Hao |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | Piccolo: A Fast and Efficient Rollback System for Virtual Machine ClustersabstractRollback is an effective technique to resume the system execution from a recorded intermediate state upon failures, without having to restart the entire system. However, in virtualized environments, rollback of a virtual machine cluster (VMC) produces high network traffic and long service disruption, particularly for a large cluster used for scientific computing, thereby imposing significant overhead both on network and applications. This paper proposes Piccolo, a fast and efficient rollback system, to restore a VMC from snapshot files over data center network. First, we exploit the similarity among VMC snapshots and leverage multicast to deliver the identical pages across VMs placed on disperse hosts, thereby bypassing unnecessary transmission of a large number of pages. Second, we analyze the impact on network traffic of varying VM placements in data center network, formulate the traffic aware placement as an optimization problem, and design a two-tier approximation algorithm that efficiently solves the problem. In addition to presenting Piccolo, we detail its implementation, and evaluate it by a set of experiments. The results show that Piccolo could achieve a significant reduction in terms of total sent data, network traffic and rollback latency compared to the existing generic techniques. Lei Cui 0003, Zhiyu Hao, Yaqiong Peng, Xiao-chun Yun |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2016 | DMNS: A Framework to Dynamically Monitor Simulated NetworkabstractWith rapid development of network simulation technology, monitoring system has become an essential tool for the researching and testing of network space activities. However, the current monitoring technologies of simulated networks cannot satisfy the requirements in terms of flexibility and efficiency. This paper proposes a framework called DMNS to dynamically monitor simulated network. With DMNS, users are able to customize the monitored objects and monitoring actions to meet the requirements of flexibility. In addition, the administrators could dynamically change the monitoring rules for saving resources based on callback mechanism. DMNS also considers the requirements of large-scale distributed simulation. Specifically, it leverages message oriented middleware to achieve efficient monitoring message information transmission to guarantee the robustness of the monitoring system. We implement a prototype of DMNS in a network range system and demonstrate the effectiveness by a case study. Zhiyu Hao, Yongzheng Zhang 0002, Yaqiong Peng, Zhenxi Sun |
ICPADS | 4 |
| 2016 | Time Donating Barrier for efficient task scheduling in competitive multicore systems
Song Wu 0001, Yaqiong Peng, Hai Jin 0001 |
Future Gener. Comput. Syst. | 2 |
| 2016 | Robinhood: Towards Efficient Work-Stealing in Virtualized EnvironmentsabstractWork-stealing, as a common user-level task scheduler for managing and scheduling tasks of multithreaded applications, suffers from inefficiency in virtualized environments, because the steal attempts of thief threads may waste CPU cycles that could be otherwise used by busy threads. This paper contributes a novel scheduling framework named Robinhood. The basic idea of Robinhood is to use the time slices of thieves to accelerate busy threads with no available tasks (referred to as poor workers) at both the guest Operating System (OS) level and Virtual Machine Monitor (VMM) level. In this way, Robinhood can reduce the cost of steal attempts and accelerate the threads doing useful work, so as to put the CPU cycles to better use. We implement Robinhood based on BWS, Linux and Xen. Our evaluation with various benchmarks demonstrates that Robinhood paves a way to efficiently run work-stealing applications in virtualized environments. Compared to Cilk++ and BWS, Robinhood can reduce up to 90 and 72 percent execution time of work-stealing applications, respectively. Yaqiong Peng, Song Wu 0001, Hai Jin 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | Towards Efficient Work-Stealing in Virtualized EnvironmentsabstractWork-stealing, as a common user-level task scheduler for managing and scheduling tasks among worker threads, has been widely adopted in multithreaded applications. With work-stealing, worker threads attempt to steal tasks from other threads' queue when they run out of their own tasks. Though work-stealing based applications can achieve good performance due to the dynamic load balancing, these steal attempting operations may frequently fail especially when available tasks are scarce, thus wasting CPU resources of busy worker threads and consequently making work-stealing less efficient in competitive environments, such as traditional multi programmed and virtualized environments. Although there are some optimizations for reducing the cost of steal-attempting threads by having such threads yield their computing resources in traditional multi programmed environments, it is more challenging to enhance the efficiency of work-stealing in virtualized environments due to the semantic gap between the virtual machine monitor (VMM) and virtual machines (VMs). In this paper, we first analyze the challenges of enhancing the efficiency of work-stealing in virtualized environments, and then propose Robin hood, a scheduling framework that reduces the cost of virtual CPUs (vCPUs) running steal-attempting threads and the scheduling delay of vCPUs running busy threads. Different from traditional scheduling methods, if the steal attempting failure occurs, Robin hood can supply the CPU time of vCPUs running steal-attempting threads to their sibling vCPUs running busy threads, which can not only improve the CPU resource utilization but also guarantee better fairness among different VMs sharing the same physical node. We implement Robin hood based on BWS, Linux and Xen. Our evaluation with various benchmarks demonstrates that Robin hood paves a way to efficiently run work-stealing based applications in virtualized platform. It can reduce up to 64% and 30% execution time of work-stealing benchmarks compared to Cilk++ and BWS respectively, and outperform credit scheduler and co-scheduling for average system throughput by 1.91× and 1.3× respectively, while guaranteeing the performance fairness among applications in virtualized environments. Yaqiong Peng, Song Wu 0001, Hai Jin 0001 |
CCGRID | 1 |
| 2015 | Optimization strategies for inter-thread synchronization overhead on NUMA machineabstractOverhead caused by data consistence issue in inter-thread synchronization probably degrades the performance of parallel applications. Non-Uniform Memory Access (NUMA), as the mainstream architecture in today's multicore processor, further exacerbates this issue due to the significant overhead incurred by Remote Memory Reference (RMR). Therefore, to reduce synchronization overhead, it is important to solve the data consistence issue. In this paper, we classify the overhead into two kinds: (1) overhead incurred by algorithms themselves, and (2) overhead incurred by critical sections. To reduce two kinds of overhead on NUMA machine, we present two optimization strategies called search and backtrace (SAB) and reorder critical section and non-critical section (RCAN), respectively. In SAB, a server thread tries to search a thread coming from master NUMA node, and designates it as the new server thread. In this way, most of the time, shared data resides in the cache of master NUMA node, resulting in lower overhead caused by data consistence issue in critical section. In RCAN, each thread consecutively posts synchronization requests, followed by consecutively executing non-critical section. In this way, server threads could serve enough requests, resulting in better data locality. We design an algorithm named R-Synch based on SAB, while designing an algorithm named H-STA based on RCAN. Our evaluation with representative synchronization algorithms demonstrates the effectiveness of R-Synch and H-STA. Song Wu 0001, Yaqiong Peng, Hai Jin 0001, Wenbin Jiang 0001 |
IPCCC | 3 |
| 2014 | Peacock: a customizable MapReduce for multicore platform
Song Wu 0001, Yaqiong Peng, Hai Jin 0001 |
J. Supercomput. | 2 |