VLDB 2026 Research / reviewers in the wild / expert
Pu Pang
dblp:172/4558
· DBLP profile ↗
18ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 4 first-author · 13 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LEGO: Supporting LLM-Enhanced Games with One Gaming GPUabstractArtificial intelligence (AI) has been increasingly applied to gaming, with large language models (LLMs) playing a key role in character control. However, efficiently co-locating game rendering and LLM inference on one GPU presents challenges due to resource constraints, diverse latency requirements, and fine-grained task scheduling. We propose LEGO, an algorithm-system co-design that enables the efficient co-location of LLM inference and game rendering tasks. Algorithmwise, LEGO features a resource-oriented layer-skipping adaptor, which distills knowledge from skipped layers to reduce computational demand while maintaining inference accuracy. System-wise, LEGO proposes a headroom-maximizing LLM scheduler, which dynamically partitions inference tasks to utilize available rendering headroom. Evaluations on an Nvidia RTX 4090 show that LEGO meets latency targets in all scenarios, improves rendering headroom utilization by up to 28.6 %, and reduces LLM inference accuracy loss by up to 86.3 % compared to current layer-skipping approaches. Han Zhao 0005, Weihao Cui, Zeshen Zhang, Jiangtong Li, Quan Chen 0002, Pu Pang, Zijun Li 0001, Zhenhua Han, Yuqing Yang 0001, Minyi Guo |
HPCA | 7 |
| 2026 | LHRS-LTW: A Load-Aware Hybrid-Cloud Resource Scheduling Framework for Large-Scale Training Workloads
Ao Wei, Jing Yang 0017, Pu Pang, Shixuan Sun, Xiaoli Ruan, Yuling Chen 0002, Minyi Guo |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2025 | WDP: Mitigating Interference in CPU Sharing Through Wake-up Delay Driven Preemption for QoS-aware Co-locationabstractAs Latency-critical (LC) tasks often experience diurnal load patterns, co-locating them with best-effort (BE) tasks improves resource utilization. Prior work allocates entire CPU cores between co-located tasks, due to the incapability of handling the interference with CPU sharing. We observed that the root cause of the interference on the same core is the inherent wake-up delay in the operating system scheduler, the wait time that a process can obtain the CPU cycles after it is woken up. Based on the finding, we propose WDP, a scheme that efficiently improves the throughput of BE tasks while ensuring QoS, leveraging CPU sharing. WDP comprises a wake-up delay-driven preemption mechanism and a preemption-based CPU manager. The preemption mechanism enables controlled preemption to reduce the wake-up delay of LC tasks with adjustable preemption capacity. Adopting the novel preemption mechanism, the CPU manager allocates CPU resources in a fine-grained manner among co-located tasks. Compared with the representative prior method, WDP improves the throughput of BE tasks by 31.2% on average while ensuring the QoS of co-located LC tasks. Yaoxuan Li, Pu Pang, Yecheng Yang, Quan Chen 0002, Zhengxuan Yan, Guoyao Xu, Liping Zhang 0013, Minyi Guo |
SoCC | 2 |
| 2025 | Generating Microservice Graphs with Production Characteristics for Efficient Resource ScalingabstractA production microservice application can have multiple services with varying call graphs, and a microservice may be shared across different call graphs.Improving resource efficiency in such complex applications requires proper benchmarks, but production traces are often too large to be used in experiments.To this end, we propose a Service Dependency Graph Generator (DGG) that comprises a Data Handler and a Graph Generator, to generate service dependency graphs of benchmarks that incorporate production-level characteristics from traces.The data handler constructs fine-grained call graphs with dynamic interface and repeated calling features from the trace, and then clusters these call graphs based on the topological and invocation types.The graph generator uses a random graph model to simulate real microservice invocations, generating multiple call graphs and merging them into small-scale service dependency graphs with production-level characteristics.Case studies show that * Fanrong Du and Jiuchen Shi contributed equally to this work. Fanrong Du, Jiuchen Shi, Quan Chen 0002, Pu Pang, Li Li 0012, Minyi Guo |
ICS | 4 |
| 2025 | ORION: Optimizing OLAP Query Execution with Proactive Caching and Separate OperatorsabstractCurrent work leverages data caching and operator execution accelerations to reduce the Online Analytical Processing (OLAP) query execution time on the disaggregated architecture with computation, cache, GPU, and storage clusters.However, their optimizations rely heavily on the OLAP engine, thus have defects of passive data fetching and integrated operator executions, leading to poor OLAP query execution performance.To resolve the above problems, we propose the ORION manager to take over the data and operator management capabilities from the OLAP engine for reducing OLAP query execution time.ORION consists of * Zhixin Tong and Jiuchen Shi contributed equally to this work. Zhixin Tong, Jiuchen Shi, Quan Chen 0002, Pu Pang, Shixuan Sun, En Shao, Minyi Guo |
ICS | 4 |
| 2025 | Reducing the End-to-End Latency of DNN-Based Recommendation Systems in GPU PoolsabstractWhile intelligent applications (e.g., recommendation systems) prefer different CPU-GPU ratios, GPU pooling technique that decouples the GPU and CPU resources yields substantial flexibility when serving diverse applications. With such architecture, DNN-based recommendation services often offload the compute-intensive neural network layers to the remote GPU pool for high resource utilization. However, such a paradigm results in the long end-to-end latency due to two causes: 1) the intermediate data is copied for multiple times during the entire process in current GPU pooling practices, incurring heavy overheads; 2) the content transferred to the GPU pool involves multiple small tensors, suffering from poor bandwidth efficiency. To solve these problems, we design Zero, a runtime system that incorporates a zero-copy transmission mechanism as well as a dynamic tensor merging policy. The zero-copy transmission mechanism unifies memory management across the inference framework and the RPC framework, accompanied by an elaborated serialization protocol to fully eliminate redundant data copying. Meanwhile, the tensor merging policy deliberately organizes small tensors into larger data blocks, so as to transfer them with higher efficiency. Experimental results show that, compared with prior work, Zero reduces the latency of typical recommendation models by up to 15.1% (10.1% on average). Guangqiang Luan, Pu Pang, Quan Chen 0002, Chen Chen 0067, Guoyao Xu, Chi Zhang 0005, Yanyi Zi, Yinghao Yu, Liping Zhang 0013, Minyi Guo |
IPDPS | 2 |
| 2025 | TSGS: Improving Gaussian Splatting for Transparent Surface Reconstruction via Normal and De-lighting PriorsabstractReconstructing transparent surfaces is essential for tasks such as robotic manipulation in labs, yet it poses a significant challenge for 3D reconstruction techniques like 3D Gaussian Splatting (3DGS). These methods often encounter a transparency-depth dilemma, where the pursuit of photorealistic rendering through standard α-blending undermines geometric precision, resulting in considerable depth estimation errors for transparent materials. To address this issue, we introduce Transparent Surface Gaussian Splatting (TSGS), a new framework that separates geometry learning from appearance refinement. In the geometry learning stage, TSGS focuses on geometry by using specular-suppressed inputs to accurately represent surfaces. In the second stage, TSGS improves visual fidelity through anisotropic specular modeling, crucially maintaining the established opacity to ensure geometric accuracy. To enhance depth inference, TSGS employs a first-surface depth extraction method. This technique uses a sliding window over α-blending weights to pinpoint the most likely surface location and calculates a robust weighted average depth. To evaluate the transparent surface reconstruction task under realistic conditions, we collect a TransLab dataset that includes complex transparent laboratory glassware. Extensive experiments on TransLab show that TSGS achieves accurate geometric reconstruction and realistic rendering of transparent objects simultaneously within the efficient 3DGS framework. Specifically, TSGS significantly surpasses current leading methods, achieving a 37.3% reduction in chamfer distance and an 8.0% improvement in F1 score compared to the top baseline. Additionally, TSGS maintains high-quality novel view synthesis, evidenced by a 0.41dB gain in PSNR, demonstrating that TSGS overcomes the transparency-depth dilemma. The code and dataset are available at https://longxiang-ai.github.io/TSGS/. Pu Pang, Hehe Fan, Hua Huang 0001, Yi Yang 0001 |
ACM Multimedia | 2 |
| 2025 | Reducing Load-Balancing Cost for Multithreading Applications on Asymmetric NUMA Machine
Yuhang Fang, Pu Pang, Quan Chen 0002, Li Li 0012, Minyi Guo |
NPC (2) | 2 |
| 2025 | EMC-LSP: A Novel Lightweight Architecture for Edge Multi-Node Long Sequence PredictionabstractEdge device traffic prediction is crucial for autonomous network control and management. However, the rapid proliferation of smart 5G networks results in increasingly heterogeneous, dynamic, and complex traffic loads on edge nodes, rendering traditional short-term prediction methods insufficient for medium- and long-term network resource scheduling. To address this, we propose a novel multi-node lightweight long-sequence deep learning-based prediction architecture (EMC-LSP) to effectively capture complex long- and short-term correlations in edge environments. Specifically, EMC-LSP employs frequency-domain hard-attention decomposition to separately model non-stationary high-frequency and low-frequency traffic, utilizes two-layer null frequency-domain convolution for long-term low-frequency similarity, and designs a high-frequency interpolation-based prediction method. In extensive tests on 18 datasets, EMC-LSP demonstrated superior performance, reducing the average prediction error of MSE and MAE by 15.20% while decreasing model parameters by 50 times. Chuanyue Xiong, Jing Yang 0017, Jiahao Zhong, Zirui He, Pu Pang, Minyi Guo |
IEEE Trans. Sustain. Comput. | 5 |
| 2023 | PMR: Priority Memory Reclaim to Improve the Performance of Latency-Critical ServicesabstractLatency-critical (LC) services are usually co-located with best-effort applications to improve the resource utilization. Lots of studies have been proposed to guarantee the performance of LC services in the co-location by managing shared resources. However, even the LC service has enough resources, its performance may still severely degrade because of memory reclaim caused by the operating system. We therefore propose Priority Memory Reclaim (PMR) which can eliminate impact of memory reclaim on LC services as much as possible. PMR consists of two techniques: priority page swapping and adaptive watermark configuration. Experiment results show that PMR can greatly improve performance of LC services. Specifically, PMR can reduce the 99%-ile latency and maximum latency of LC services by up to 92.32% and 95.87% while improving the throughput by up to 4.75x. Bo Liu 0122, Kaihao Bai, Pu Pang, Quan Chen 0002, Yaoxuan Li, Minyi Guo |
ICPADS | 3 |
| 2023 | PAC: Preference-Aware Co-location Scheduling on Heterogeneous NUMA Architectures To Improve Resource UtilizationabstractLatency-critical applications directly interact with end users and often experience the diurnal load pattern. In production, best-effort applications are often co-located with them to utilize the idle cores at the low load. Meanwhile, modern computers are evolving towards heterogeneous NUMA architecture, where the cores have different computation abilities, memory access latencies and network communication delays. Prior co-location scheduling work did not consider the NUMA architecture, and failed to maximize the throughput of best-effort applications while ensuring the required QoS of latency-critical applications. Our investigation shows that NUMA effect has complex impacts on the latency of latency-critical applications and the throughput of best-effort applications. We therefore propose PAC, a preference-aware co-location scheduling scheme that considers the NUMA effect for heterogeneous NUMA architectures. PAC has a performance monitor and a core scheduler. Specifically, the performance monitor identifies the "dangerous" latency-critical applications that require upgrading core allocations. We propose two low-overhead scheduling strategies for the scheduler. The strategies identify the bottlenecks of applications and adjust core allocations accordingly. Experimental result shows that PAC improves the throughput of best-effort applications by 3.87× while ensuring the required QoS of latency-critical applications. Pu Pang, Yaoxuan Li, Bo Liu 0122, Quan Chen 0002, Zhou Yu 0003, Zhibin Yu 0001, Deze Zeng, Jingwen Leng, Jieru Zhao, Minyi Guo |
ICS | 1 |
| 2023 | Async-fork: Mitigating Query Latency Spikes Incurred by the Fork-based Snapshot Mechanism from the OS LevelabstractIn-memory key-value stores (IMKVSes) serve many online applications. They generally adopt the fork-based snapshot mechanism to support data backup. However, this method can result in query latency spikes because the engine is out-of-service for queries during the snapshot. In contrast to existing research optimizing snapshot algorithms, we address the problem from the operating system (OS) level, while keeping the data persistent mechanism in IMKVSes unchanged. Specifically, we first study the impact of the fork operation on query latency. Based on findings in the study, we propose Async-fork, which performs the fork operation asynchronously to reduce the out-of-service time of the engine. Async-fork is implemented in the Linux kernel and deployed into the online Redis database in public clouds. Our experiment results show that Async-fork can significantly reduce the tail latency of queries during the snapshot. Pu Pang, Kaihao Bai, Quan Chen 0002, Shixuan Sun, Bo Liu 0122, Hongbo Yao, Zhengheng Wang, Zheng Liu 0022, Yong Yang 0013, Tao Ma 0006, Minyi Guo |
Proc. VLDB Endow. | 1 |
| 2022 | CSC: Collaborative System Configuration for I/O-Intensive Applications in Multi-Tenant CloudsabstractI/O-intensive applications are important workloads of public clouds. Multiple cloud applications co-run on the same physical machine in different virtual machines (VMs), and the shared resources (e.g., disk bandwidth) are often isolated for fairness. Our investigation shows that the performance of an I/O-intensive application is impacted by both disk bandwidth allocation and the page cache settings in the guest operating system. However, none of prior work considers adjusting the page cache settings for better performance, when the disk bandwidth allocation is adjusted. We therefore propose CSC, a system that collaboratively identifies the appropriate disk bandwidth allocation and page cache settings in the guest operating system of each VM. CSC aims to improve the system-wide I/O throughput of the physical machine, while also improve the I/O throughput of each individual I/O-intensive application in VMs. CSC comprises an online disk bandwidth allocator and an adaptive dirty page setting optimizer. The bandwidth allocator monitors the disk bandwidth utilization and re-allocates some bandwidth from free VMs to busy VMs periodically. After the re-allocation, the opti-mizer identifies the appropriate dirty page settings in the guest operating system of the VMs using Bayesian Optimization. The experimental results show that CSC improves the performance of I/O-intensive applications by 9.5 % on average (up to 17.29 %) when 5 VMs are co-located while fairness is guaranteed. Haowei Huang, Pu Pang, Quan Chen 0002, Jieru Zhao, Wenli Zheng, Minyi Guo |
IPDPS | 2 |
| 2022 | Online Thread Auto-Tuning for Performance Improvement and Resource SavingabstractMulti-threading is a common way for programs to benefit from the multi/many-core design. However, the performance of some parallel programs does not increase/even decrease as the number of cores/threads increases. Our study shows that the performance of a parallel program is impacted bythe number of cores/threads,the thread placement,the inputs of the program. It is nontrivial to identify the optimal number of cores and the corresponding thread placement to maximize the performance, when the input of a program is determined online and the workload of different iterations may not be identical. To resolve the above problem, we proposeOtter, a thread auto-tuning system at runtime for iterative parallel programs. Otter collects the runtime information in the first few iterations and makes decisions on the number of threads and thread placement policy to achieve the goal of improving performance or saving resources. It considers the characteristics of dynamic workload in the iteration process and reduces the time overhead through a migration method. Experiments on a 96-core machine show that Otter improves the performance of the benchmarks by 20.7% and reduces core hours by 51.3% on average compared to the case of running them with all the CPU cores. Guangqiang Luan, Pu Pang, Quan Chen 0002, Minyi Guo |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | Adaptive Preference-Aware Co-Location for Improving Resource Utilization of Power Constrained DatacentersabstractLarge-scale datacenters often host latency-sensitive services that have stringent Quality-of-Service requirement and experience diurnal load pattern. Co-locating best-effort applications that have no QoS requirement with the latency-sensitive services has been widely used to improve the resource utilization of datacenters with careful shared resource management. However, existing co-location techniques tend to result in the power overload problem on power constrained servers due to the ignorance of the power consumption. To this end, we propose Sturgeon, a runtime system proactively manages resources between co-located applications in a power constrained environment, to ensure the QoS of latency-sensitive services while maximizing the throughput of best-effort applications. Our investigation shows that, at a given load, there are multiple feasible resource configurations to meet both QoS requirement and power budget, while one of them yields the maximum throughput of best-effort applications. To find such a configuration, we establish models to accurately predict the performance and power consumption of the co-located applications. Sturgeon monitors the QoS of the services periodically, in order to eliminate the potential QoS violation caused by the unpredictable interference. Besides, when the datacenter hosts different types of applications to perform co-location, Sturgeon places applications with their preferable candidates to improve the overall throughput. The experimental results show that at server level Sturgeon improves the throughput of the best-effort application by 25.43 percent compared to the state-of-the-art technique, while guaranteeing the 95%-ile latency within the QoS target; at cluster level, Sturgeon improves the overall throughput of best-effort applications by 13.74 percent compared to the baseline. Pu Pang, Quan Chen 0002, Deze Zeng, Minyi Guo |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | Sturgeon: Preference-aware Co-location for Improving Utilization of Power Constrained ComputersabstractLarge-scale datacenters often host latency-sensitive services that have stringent Quality-of-Service requirement and experience diurnal load pattern. Co-locating best-effort applications that have no QoS requirement with latency-sensitive services has been widely used to improve the resource utilization with careful shared resource management. However, existing co-location techniques tend to result in the power overload problem on power constrained computers due to the ignorance of the power consumption. To this end, we propose Sturgeon, a runtime system proactively manages resources between colocated applications in a power constrained environment, to ensure the QoS of latency-sensitive services while maximizing the resource utilization. Our investigation shows that, at a given load, there are multiple feasible resource configurations to meet both QoS requirement and power budget, while one of them yields the maximum throughput of best-effort applications. To find such a configuration, we establish models to accurately predict the performance and power consumption of the colocated applications. Sturgeon monitors the QoS periodically in order to eliminate the potential QoS violation caused by the unpredictable interference. The experimental results show that Sturgeon improves the throughput of best-effort applications by 24.96% compared to the state-of-the-art technique, while guaranteeing the 95%-ile latency within the QoS target. Pu Pang, Quan Chen 0002, Deze Zeng, Chao Li 0009, Jingwen Leng, Wenli Zheng, Minyi Guo |
IPDPS | 1 |
| 2018 | In-growth test for monolithic 3D integrated SRAMabstractMonolithic three-dimensional integration (M3I) directly fabricates tiers of integrated circuits upon each other and provides millions of vertical interconnections with inter-layer vias (ILVs). It thus brings higher integration density and communication capability compared with three-dimensional stacked integration (3D-SI). However, the Known-Good-Die problem haunting 3D-SI-a faulty tier causes the failure of the entire stack-also occurs in M3I. Lack of efficient test methodologies such as the pre-bond testing in 3D-SI, M3I may have a more significant yield drop and thus its cost may be unacceptable for main-stream adoption. This paper introduces a novel In-growth test method for M3I SRAM. We propose a novel Design-for-Test (DfT) methodology to enable the proposed In-growth test on cell-level partitioned incomplete SRAM cells. We also build a statistical model of cost and discover a prospective judgement to determine whether or not to stop the fabrication, in order to prevent from raising the cost of fabricating more tiers upon the irreparable tiers. We find that a “sweet point” exists in the judgement, which can minimize the overall cost. Experimental results show the effectiveness of our proposed test methodology. Pu Pang, Yixun Zhang, Tianjian Li, Sung Kyu Lim, Quan Chen 0002, Xiaoyao Liang, Li Jiang 0002 |
DATE | 1 |
| 2015 | On diagnosable and tunable 3D clock network design for lifetime reliability enhancementabstractIn three-dimensional (3D) integrated circuits (IC-s), many clock-TSVs are deployed to deliver clock signals to different tiers with minimum skews. However, these clock-TSVs are prone to aging effects, such as thermal-mechanical stress and electromigration, rendering hard-to-predict clock skews at runtime. These skews have a wide range of influence on the flip-flops, and may violate the safety margins of critical paths in the circuit. Besides the circuit aging effect, the clock-TSV induced skews pose another threat to the circuit lifetime reliability. To tackle this problem, we propose to put tunable buffer for each clock-TSV in the clock network, and introduce an efficient algorithm to place aging sensors in the circuit at design stage. Then, at runtime, we conduct online diagnosis and apply effective clock tuning algorithms based on the triggered alarms in the aging sensors. Experimental results on a post-layout 3D circuit show that the proposed solution is able to significantly improve the lifetime reliability of 3D ICs. Li Jiang 0002, Pu Pang, Naifeng Jing, Sung Kyu Lim, Xiaoyao Liang, Qiang Xu 0001 |
ITC | 2 |