VLDB 2026 Research / reviewers in the wild / expert
Yunhao Yao
dblp:343/8898
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-7433-3145ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 4 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GCA-BULF: A Bottom-Up Framework for Short-Term Load Forecasting Using Grouped Critical AppliancesabstractWith the rise of time-of-use and tiered electricity pricing, energy consumers are encouraged to adopt peak-shifting strategies by automatically controlling high-power appliances. These help lower energy costs while enhancing the power grid's stability. To support such energy management with high resilience and responsiveness, reliable short-term load forecasting (STLF) plays a critical role. STLF predicts electricity consumption over time horizons ranging from minutes to days, using historical data, temporal patterns, and contextual factors. Traditional top-down forecasting methods struggle to capture the complex consumption patterns of diverse and mixed appliance loads. Although bottom-up methods improve forecasting accuracy by integrating appliance-level data, monitoring all appliances is costly, and many do not meaningfully impact total load prediction. Therefore, we propose GCA-BULF, a bottom-up short-term load forecasting framework based on grouped critical appliances, supported by three key designs. First, the Critical Appliance Filtering module ranks appliances according to their power consumption, switching frequency, and usage pattern periodicity, and identifies critical ones through iterative load decomposition. Next, the Related Appliance Grouping module clusters these appliances based on spatial and temporal correlations for group-level forecasting. Finally, the Collaborative Load Forecasting module refines the total load prediction by combining multiple group-level forecasts. We evaluate GCA-BULF on residential and office building load forecasting tasks. Experimental results reveal that GCA-BULF improves hourly total load forecasting by 20.85%-57.88% compared to existing top-down methods and by 33.03%-92.48% compared to bottom-up methods. Yunhao Yao, Jinwei Fang, Puhan Luo, Jiahui Hou, Xiang-Yang Li 0001 |
IWQoS | 1 |
| 2026 | Beyond Detection: Autonomous Anomaly Remediation for MCP Against Tool Poisoning AttacksabstractLLM-powered agents are evolving from passive recommenders into autonomous executors, leveraging tools via the Model Context Protocol (MCP) for web automation. However, this paradigm introduces a new vulnerability: tool poisoning attacks that manipulate the MCP context can corrupt an agent's reasoning. Existing methods focus on anomaly detection and lack autonomous correction mechanisms, hindering their real-world deployment. Guanquan Shi, Yichao Gao, Hongsen Lang, Yunhao Yao, Haohua Du, Xiang-Yang Li 0001 |
WWW | 6 |
| 2026 | Gproxy: Communication-Efficient Federated Graph Learning With Efficient Adaptive ProxyingabstractFederated graph learning (FGL) enables multiple participants with distributed but connected graph data to collaboratively train a model in a privacy-preserving way. However, the high communication cost hinders the adoption of FGL in many resource-limited or delay-sensitive applications. In this work, we focus on reducing the communication cost incurred by the transmission of neighborhood information in FGL. We propose to search for local proxies that can play a substitute role as the external neighbors and develop a novel federated graph learning framework namedGproxy.Gproxyutilizes representation similarity and class correlation to select local proxies for external neighbors. Additionally, we propose to dynamically adjust the proxy strategy according to the changing representation of nodes during the iterative training process. We also design a proxy cache to accelerate the search process by reusing proxy search outcomes for similar external neighbors. Furthermore, we provide a theoretical analysis and show that using a proxy node has a similar influence on training when it is sufficiently similar to the external one. Extensive evaluations show thatGproxysignificantly reduces communication cost while maintaining model performance compared to strong baselines. Junyang Wang 0004, Lan Zhang 0002, Mu Yuan, Yihang Cheng 0002, Yunhao Yao, Zhonghao Hu |
IEEE Trans. Mob. Comput. | 5 |
| 2026 | PrivGuardInfer: Channel-Level End-Edge Collaborative Inference Strategy Protecting Original Inputs and Sensitive AttributesabstractEnd-edge collaborative inference improves computational efficiency by dividing a deep neural network into two parts, executed across the end device and the edge node in parallel. However, adversaries like malicious edge nodes can exploit transmitted data to reconstruct original inputs or infer sensitive attributes. Existing collaborative inference strategies upload the majority of input features to the edge node, significantly increasing the risk of privacy leakage, even without input reconstruction. Therefore, we propose PrivGuardInfer, a channel-level DNN end-edge collaborative inference strategy that optimizes intra-layer partition to simultaneously protect original inputs and sensitive attributes while ensuring latency constraints, supported by three key designs. First, the privacy measurements oriented both layer depth and channel count, jointly quantify the difficulty of reconstructing original inputs using varying numbers of feature maps across different layers. After assessing each channel's contribution, the information offset further measures the difficulty of inferring sensitive attributes. Finally, PrivGuardInfer models the privacy-optimal intra-layer partition under latency constraints as a grouped knapsack problem, mapping attack difficulty to item values and inference latency to item weights. Experimental results reveal that PrivGuardInfer achieves an average improvement of 80.54% in defending against model inversion attacks and 63.34% against attribute inference attacks compared to existing end-edge partition strategies. Moreover, it outperforms current privacy protection methods by an average of 69.37% and 49.75% in mitigating these two types of attacks. Yunhao Yao, Puhan Luo, Yihang Cheng 0002, Jiahui Hou, Xiang-Yang Li 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | MDRPASS: A Multi-Dimensional Demand Response Potential Assessment Based Scheduling Strategy for Smart GridabstractThe high penetration of renewable energy presents significant challenges for the power grid in balancing supply and demand. Demand response (DR) is crucial for stable grid operation, yet existing research often overlooks actual electricity consumption patterns for the following day and the diversity among user types, compromising assessment accuracy and applicability. To address these shortcomings, this paper proposes an integrated load regulation system comprising “Identification-Prediction-Scheduling”. First, we identify specific user groups and utilize load prediction models to accurately forecast future loads. By analyzing predicted loads alongside historical data, we uncover electricity consumption patterns and DR potential for the following day, providing better insights for power grid scheduling. Finally, we implement a load control strategy aimed at optimizing grid stability. Our results show that the load forecasting model achieves average errors of 17.86% and 16.16% at intervals of 15 minutes and 1 hour, respectively. When demand is set at 1500 kW, the proposed load control strategy predicts a response load curtailment of 2216.85 kW and a DPI score of 50.75, significantly outperforming both the random selection strategy (1887.11 kW, DPI score of 18.39) and the priority to large FBC strategy (288.86 kW, DPI score of 40.78). This approach ensures better enterprise selection, contractual compliance, and improved benefits for electrical users. Siyu Jing, Yunhao Yao, Haishi Du, Jinwei Fang, Jiahui Hou, Xiang-Yang Li 0001 |
ICC | 2 |
| 2025 | Task-Oriented Training Data Privacy Protection for Cloud-based Model Training
Jiahui Hou, Haifeng Sun 0005, Jingmiao Zhang, Yunhao Yao, Haikuo Yu, Xiang-Yang Li 0001 |
USENIX Security Symposium | 5 |
| 2025 | TrafficDiary: User Attribute Inference Based on Smart Home Traffic TracesabstractSmart home technology has found wide-ranging applications in daily life, from enhancing energy efficiency to simplifying daily tasks and providing greater convenience. However, recent works have found that smart home devices are vulnerable to passive network observers (i.e., adversaries). While adversaries have demonstrated the ability to infer device events (e.g., whether a lamp is turned on) from the encrypted smart home traffic, we believe this only represents a less critical aspect of smart home privacy risks. Further analysis of demographic attributes presents greater risks to user privacy. Besides, from our deployment experience of real-world smart homes, we found that existing event inference methods can be greatly interfered with by event-unrelated traffic. Experiments show that this interference can result in up to a 10% drop in inference accuracy. Furthermore, it is challenging to infer finer-grained demographic attributes, due to the insufficient accuracy of event inference. Therefore, in this work, we propose a novel event inference model extracting multi-dimensional features that reduces the interference of event-unrelated traffic by analyzing packet length distribution and statistical properties. In addition, we design a dual-channel neural network to extract spatial and temporal relationships among triggered events to infer demographic attributes of smart home users, such as age group and career stage. Combining the above designs, we present TrafficDiary, the first user attribute inference approach based on smart home traffic traces. We prototype TrafficDiary and evaluate it in real-world smart homes. Experimental results show that TrafficDiary achieves 98.68% accuracy with a zero false positive rate in event inference and a high level of accuracy in user attribute inference, even when 16, 362 groups of event-unrelated traffic exist. TrafficDiary also performs well in terms of efficiency, with an inference latency of only 1.82 ms on a Raspberry Pi 4B device. Yunhao Yao, Jiahui Hou, Mu Yuan, Zhengyuan Xu, Xiang-Yang Li 0001 |
ACM Trans. Internet Techn. | 1 |
| 2024 | F2Zip: Finetuning-Free Model Compression for Scenario-Adaptive Embedded VisionabstractWith the development of the Internet of Things and artificial intelligence, the deployment and inference of intelligent models have gradually raised concerns. To reduce the huge computation and storage overhead of modern deep neural networks, many studies use model pruning techniques to reduce the model size and computational cost. However, existing pruning techniques usually require model fine-tuning, which incurs high additional overhead, making them difficult to apply to real-world scenarios. In this work, we focus on vision model compression and present F2Zip, a scenario-adaptive finetuning-free pruning framework for embedded devices. First, we propose a scenario complexity measurement that quantifies scenario changes with pixel-level entropy. By analyzing the scenario complexity, F2Zip adaptively evaluates the importance of different channels and layers of the model using only a small amount (tens) of unlabeled data. Then we design a multi-constraint knapsack solver to prune scenario-unrelated redundant channels. We implemented and deployed F2Zip in surveillance scenarios and tested different models on videos collected from both public and real-world sources. Experimental results show that F2Zip is free of model fine-tuning in various scenarios. F2Zip reduces the end-to-end deployment time by 89.8% and reduces energy cost by 79.5%, which shows that F2Zip is computationally friendly for embedded devices. Without fine-tuning and any accuracy degradation, F2Zip achieves up to 50.2% parameter reduction, outperforming baseline methods by 35.1%. Puhan Luo, Jiahui Hou, Mu Yuan, Yunhao Yao, Xiang-Yang Li 0001 |
SenSys | 5 |
| 2024 | SecoInfer: Secure DNN End-Edge Collaborative Inference Framework Optimizing Privacy and LatencyabstractEnd-edge collaborative inference enhances computational efficiency by segmenting a deep neural network (DNN) model into two parts, executed across the end device and the edge node. However, existing collaborative inference strategies often involve transmitting original inputs from the end device to the edge node, resulting in significant risks of user detail leakage without requiring input reconstruction. Therefore, in this work, we present SecoInfer, a secure layer-level DNN end-edge collaborative inference framework. SecoInfer achieves joint optimization of data privacy and inference latency for DNN partition solutions that meet latency constraints, supported by three key designs. First, the privacy-aware DNN layer projection measurement quantifies the difficulty adversaries encounter in reconstructing the original input from the intermediate output of each layer. Then, the latency-privacy integrated structure modeling enables the direct calculation of the privacy measurement and inference latency for each partition solution from a list element or a directed acyclic graph (DAG) cut. Finally, the two-stage latency constraint adjustment scheme narrows down the search space of feasible partition solutions at the block level and fine-tunes the final one to meet the latency constraint based on layer depth. We prototype SecoInfer, utilizing a Raspberry Pi 4B as the end device and a server with an NVIDIA GeForce RTX 3060 GPU as the edge node. Experimental results demonstrate that under latency constraints of 20 ms, 33 ms, and 40 ms, SecoInfer reduces adversarial data reconstruction by 9.84%, 19.26%, and 25.18%, respectively, without any loss of task model accuracy. SecoInfer also enhances efficiency, reducing the time needed to determine optimal end-edge partition solutions on a Raspberry Pi 4B by 18.04%. Yunhao Yao, Jiahui Hou, Yihang Cheng 0002, Mu Yuan, Puhan Luo, Xiang-Yang Li 0001 |
ACM Trans. Sens. Networks | 1 |
| 2022 | Traffic Processing and Fingerprint Generation for Smart Home Device EventabstractRecent studies show that smart home devices are vulnerable to passive network observers, referred to as adversaries. An adversary can use mobile devices as a sniffer to infer device events (e.g., a door is opened/closed) even through encrypted WiFi traffic. However, to obtain high event identification accuracy, existing works heavily rely on either machine learning (ML) or large traffic feature sizes, which results in a costly retraining or pairing process. In this paper, we generate an event traffic fingerprint and propose an event identification method with good extensibility. To address the interference of network fluctuations and unrelated traffic, we filter all redundant and event-unrelated packets and extract event packet sequences based on time interval. After processing, we generate a representative packet-level fingerprint for each event and identify the event based on fingerprint matching. The precision and recall of event identification reach 96.77% and 92.31% on average. Our traffic processing method can work as an add-on to existing ML-based event inference methods, which leads to at least 12% increase on the accuracy. Yunhao Yao, Jiahui Hou, Zhengyuan Xu, Xiang-Yang Li 0001 |
ICPADS | 1 |