VLDB 2026 Research / reviewers in the wild / expert
Puhan Luo
dblp:322/7804
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
0009-0007-6885-1185ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GCA-BULF: A Bottom-Up Framework for Short-Term Load Forecasting Using Grouped Critical AppliancesabstractWith the rise of time-of-use and tiered electricity pricing, energy consumers are encouraged to adopt peak-shifting strategies by automatically controlling high-power appliances. These help lower energy costs while enhancing the power grid's stability. To support such energy management with high resilience and responsiveness, reliable short-term load forecasting (STLF) plays a critical role. STLF predicts electricity consumption over time horizons ranging from minutes to days, using historical data, temporal patterns, and contextual factors. Traditional top-down forecasting methods struggle to capture the complex consumption patterns of diverse and mixed appliance loads. Although bottom-up methods improve forecasting accuracy by integrating appliance-level data, monitoring all appliances is costly, and many do not meaningfully impact total load prediction. Therefore, we propose GCA-BULF, a bottom-up short-term load forecasting framework based on grouped critical appliances, supported by three key designs. First, the Critical Appliance Filtering module ranks appliances according to their power consumption, switching frequency, and usage pattern periodicity, and identifies critical ones through iterative load decomposition. Next, the Related Appliance Grouping module clusters these appliances based on spatial and temporal correlations for group-level forecasting. Finally, the Collaborative Load Forecasting module refines the total load prediction by combining multiple group-level forecasts. We evaluate GCA-BULF on residential and office building load forecasting tasks. Experimental results reveal that GCA-BULF improves hourly total load forecasting by 20.85%-57.88% compared to existing top-down methods and by 33.03%-92.48% compared to bottom-up methods. Yunhao Yao, Jinwei Fang, Puhan Luo, Jiahui Hou, Xiang-Yang Li 0001 |
IWQoS | 3 |
| 2026 | PrivGuardInfer: Channel-Level End-Edge Collaborative Inference Strategy Protecting Original Inputs and Sensitive AttributesabstractEnd-edge collaborative inference improves computational efficiency by dividing a deep neural network into two parts, executed across the end device and the edge node in parallel. However, adversaries like malicious edge nodes can exploit transmitted data to reconstruct original inputs or infer sensitive attributes. Existing collaborative inference strategies upload the majority of input features to the edge node, significantly increasing the risk of privacy leakage, even without input reconstruction. Therefore, we propose PrivGuardInfer, a channel-level DNN end-edge collaborative inference strategy that optimizes intra-layer partition to simultaneously protect original inputs and sensitive attributes while ensuring latency constraints, supported by three key designs. First, the privacy measurements oriented both layer depth and channel count, jointly quantify the difficulty of reconstructing original inputs using varying numbers of feature maps across different layers. After assessing each channel's contribution, the information offset further measures the difficulty of inferring sensitive attributes. Finally, PrivGuardInfer models the privacy-optimal intra-layer partition under latency constraints as a grouped knapsack problem, mapping attack difficulty to item values and inference latency to item weights. Experimental results reveal that PrivGuardInfer achieves an average improvement of 80.54% in defending against model inversion attacks and 63.34% against attribute inference attacks compared to existing end-edge partition strategies. Moreover, it outperforms current privacy protection methods by an average of 69.37% and 49.75% in mitigating these two types of attacks. Yunhao Yao, Puhan Luo, Yihang Cheng 0002, Jiahui Hou, Xiang-Yang Li 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | A-VL: Adaptive Attention for Large Vision-Language ModelsabstractThe Large Vision-Language Model (LVLM) integrates computer vision and natural language processing techniques, offering substantial application potential. However, these models demand extensive resources during inference. Adaptive attention techniques can dynamically reduce computational redundancy and thus improve efficiency. Although current adaptive attention methods significantly reduce the memory requirements of Transformer-based language models, they are not tailored for LVLMs. We observe that LVLMs generate responses from both remote image tokens and local text tokens, and different modalities have different attention patterns. This observation inspires us to manage the attention for each modality separately. Specifically, for visual input, we store the cache of potentially useful information but only compute the most critical parts. For language input, we care more about local information. Based on our observation and analysis of vision-language attention patterns, we develop A-VL, a plug-and-play adaptive attention tailored for LVLM inference. Extensive evaluations on three vision-language tasks and five datasets show the effectiveness of our designs. Our approach A-VL outperforms existing adaptive attention methods in reducing memory usage and computational load without compromising performance. Junyang Zhang 0001, Mu Yuan, Ruiguang Zhong, Puhan Luo, Huiyou Zhan, Ningkang Zhang, Chengchen Hu, Xiang-Yang Li 0001 |
AAAI | 4 |
| 2025 | ChannelZip: SLO-Aware Channel Compression for Task-Adaptive Model Serving on IoT DevicesabstractDeploying deep neural networks (DNNs) on IoT devices for model serving is a promising solution for intelligent applications with high real-time requirements and bandwidth sensitivity. To cope with the prohibitive computation and storage overheads of modern DNNs, great efforts have been devoted to the model compression technique. Most existing model compression approaches focus on minimizing the model size and maximizing the average accuracy on all the inference tasks. However, real-world IoT tasks have various service-level objectives (SLOs). Models compressed by existing methods struggle to simultaneously meet SLOs in multiple dimensions, such as latency and accuracy. In this work, we study model compression with a joint consideration of SLO awareness and task adaptation. Through our extensive experience with model compression across various IoT tasks, we observe that the importance of individual channels in contributing to accuracy is heavily influenced by task-specific data distribution. Therefore, we design a channel Shapley algorithm to estimate the importance of individual channels in DNNs and propose a deep reinforcement learning based controller to incorporate SLOs into the compression objective. Integrating these designs, we propose and prototype ChannelZip, the first SLO-aware channel compression framework. Extensive evaluations on real IoT model serving systems show the effectiveness in task adaptation of ChannelZip. ChannelZip outperforms strong model compression baselines by 3.77% accuracy and achieves a 69% average parameter compression ratio. Real-world deployment on different IoT devices shows that ChannelZip meets all task SLOs and achieves up to 2.32 × inference speedup. Puhan Luo, Jiahui Hou, Haisheng Tan, Mu Yuan, Xiang-Yang Li 0001 |
ACM Trans. Sens. Networks | 1 |
| 2024 | F2Zip: Finetuning-Free Model Compression for Scenario-Adaptive Embedded VisionabstractWith the development of the Internet of Things and artificial intelligence, the deployment and inference of intelligent models have gradually raised concerns. To reduce the huge computation and storage overhead of modern deep neural networks, many studies use model pruning techniques to reduce the model size and computational cost. However, existing pruning techniques usually require model fine-tuning, which incurs high additional overhead, making them difficult to apply to real-world scenarios. In this work, we focus on vision model compression and present F2Zip, a scenario-adaptive finetuning-free pruning framework for embedded devices. First, we propose a scenario complexity measurement that quantifies scenario changes with pixel-level entropy. By analyzing the scenario complexity, F2Zip adaptively evaluates the importance of different channels and layers of the model using only a small amount (tens) of unlabeled data. Then we design a multi-constraint knapsack solver to prune scenario-unrelated redundant channels. We implemented and deployed F2Zip in surveillance scenarios and tested different models on videos collected from both public and real-world sources. Experimental results show that F2Zip is free of model fine-tuning in various scenarios. F2Zip reduces the end-to-end deployment time by 89.8% and reduces energy cost by 79.5%, which shows that F2Zip is computationally friendly for embedded devices. Without fine-tuning and any accuracy degradation, F2Zip achieves up to 50.2% parameter reduction, outperforming baseline methods by 35.1%. Puhan Luo, Jiahui Hou, Mu Yuan, Yunhao Yao, Xiang-Yang Li 0001 |
SenSys | 1 |
| 2024 | QoS-Ensured Model Optimization for AIoT: A Multi-Scale Reinforcement Learning ApproachabstractOptimizing deep neural network (DNN) models to meet quality of service (QoS) requirements in terms of accuracy and computation is of crucial importance for realizing efficient on-device inference in resource-constrained artificial intelligence of things (AIoT). However, most existing works can hardly satisfy the aforementioned QoS requirements since the intrinsic multi-scale characteristic of DNN structures has been seldom considered. In this paper, we formulate a QoS-ensured DNN model structure optimization problem as a novel multi-scale Markov decision process (MSMDP), which can collaboratively decide the DNN structures from different scales. To efficiently solve the above problem, we propose a multi-scale reinforcement learning (MSRL) algorithm, which jointly optimizes block and channel number by interactive multi-scale decision, while ensuring QoS by QoS-based decision evaluation and policy update. Extensive experiments are conducted in both the actual AIoT scenarios and public datasets for different tasks by using different AIoT devices. The results confirm that our proposed MSRL outperforms the baseline schemes in terms of QoS satisfaction, convergence performance, and complexity. Specifically, our algorithm respectively reduces 98.6% computation and 95.7% model size at most while ensuring the QoS compared with the state-of-the-art methods. Fuhui Zhou, Yuben Qu, Puhan Luo, Xiang-Yang Li 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Efficient Task-driven Video Data Privacy Protection for Smart Camera Surveillance SystemabstractAs one of the most commonly used AIoT sensors, smart cameras and their supporting services, namely cloud video surveillance (CVS) systems, have brought great convenience to people’s lives. Recent CVS providers use different machine learning techniques to improve their services (regarded as tasks) based on the uploaded video. However, uploading data to the CVS providers may cause severe privacy issues. Existing works that remove privacy information could not achieve a high tradeoff between data usability and privacy, because the importance of information varies with the task. In addition, it is challenging to design a real-time privacy protection mechanism, especially in resource-constrained smart cameras. In this work, we design a task-driven and efficient video privacy protection mechanism for a better tradeoff between privacy and data usability. We use Class Activation Mapping to protect privacy while preserving data usability. To improve the efficiency, we utilize the motion vector and residual matrix produced during video codec. Our work outperforms the region of interest–based methods in data protection while preserving data usability. The attack accuracy drops 70%, while the task accuracy is comparable to those without protection (within ± 4%). The average protection frame rate of the High Definition video can exceed 16 fps+ even on a CPU. Jiahui Hou, Suyuan Liu, Puhan Luo, Xiang-Yang Li 0001 |
ACM Trans. Sens. Networks | 5 |
| 2024 | SecoInfer: Secure DNN End-Edge Collaborative Inference Framework Optimizing Privacy and LatencyabstractEnd-edge collaborative inference enhances computational efficiency by segmenting a deep neural network (DNN) model into two parts, executed across the end device and the edge node. However, existing collaborative inference strategies often involve transmitting original inputs from the end device to the edge node, resulting in significant risks of user detail leakage without requiring input reconstruction. Therefore, in this work, we present SecoInfer, a secure layer-level DNN end-edge collaborative inference framework. SecoInfer achieves joint optimization of data privacy and inference latency for DNN partition solutions that meet latency constraints, supported by three key designs. First, the privacy-aware DNN layer projection measurement quantifies the difficulty adversaries encounter in reconstructing the original input from the intermediate output of each layer. Then, the latency-privacy integrated structure modeling enables the direct calculation of the privacy measurement and inference latency for each partition solution from a list element or a directed acyclic graph (DAG) cut. Finally, the two-stage latency constraint adjustment scheme narrows down the search space of feasible partition solutions at the block level and fine-tunes the final one to meet the latency constraint based on layer depth. We prototype SecoInfer, utilizing a Raspberry Pi 4B as the end device and a server with an NVIDIA GeForce RTX 3060 GPU as the edge node. Experimental results demonstrate that under latency constraints of 20 ms, 33 ms, and 40 ms, SecoInfer reduces adversarial data reconstruction by 9.84%, 19.26%, and 25.18%, respectively, without any loss of task model accuracy. SecoInfer also enhances efficiency, reducing the time needed to determine optimal end-edge partition solutions on a Raspberry Pi 4B by 18.04%. Yunhao Yao, Jiahui Hou, Yihang Cheng 0002, Mu Yuan, Puhan Luo, Xiang-Yang Li 0001 |
ACM Trans. Sens. Networks | 6 |
| 2023 | PianoWatch: An Intelligent Piano Understanding and Evaluation System Using SmartwatchabstractExisting intelligent piano learning systems mainly assist the player by camera, which generally only consider the fingering that can only reflect the performance problem from a limited perspective. Thus, we propose PianoWatch, a multi-dimensional assistance system based on wrist wearable devices. PianoWatch extracts more accurate patterns by analyzing data from microphone, camera, accelerometers and gyroscopes. Then it gives corresponding playing advices, in addition to pitch, including fingering, depth and even mood, through an evaluation model. We implement the prototype of PianoWatch and evaluate it by 20 volunteers. Extensive experiments show it can achieve 93% F1-Score on the playing pattern extraction, and more than 95% of users would like to try it to assist their piano learning. Hao Zhou 0001, Siyu Jing, Haohua Du, Puhan Luo, Jiahui Hou, Xiang-Yang Li 0001 |
SECON | 5 |