VLDB 2026 Research / reviewers in the wild / expert
Yihang Cheng 0002
dblp:184/7176-2
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-9670-0196ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TopFGL: A Topology-Aware and Distributionagnostic Federated Learning Framework Tackling Topological Heterogeneity on Graph Data
Junyang Wang 0004, Lan Zhang 0002, Yihang Cheng 0002, Mu Yuan, Tian Wang 0001, Zhihui Fu |
ICDE | 3 |
| 2026 | STIP: Three-Party Privacy-Preserving and Lossless Inference for Large Transformers in Production
Mu Yuan, Lan Zhang 0002, Yihang Cheng 0002, Miaohui Song, Guoliang Xing, Xiang-Yang Li 0001 |
NDSS | 3 |
| 2026 | Gproxy: Communication-Efficient Federated Graph Learning With Efficient Adaptive ProxyingabstractFederated graph learning (FGL) enables multiple participants with distributed but connected graph data to collaboratively train a model in a privacy-preserving way. However, the high communication cost hinders the adoption of FGL in many resource-limited or delay-sensitive applications. In this work, we focus on reducing the communication cost incurred by the transmission of neighborhood information in FGL. We propose to search for local proxies that can play a substitute role as the external neighbors and develop a novel federated graph learning framework namedGproxy.Gproxyutilizes representation similarity and class correlation to select local proxies for external neighbors. Additionally, we propose to dynamically adjust the proxy strategy according to the changing representation of nodes during the iterative training process. We also design a proxy cache to accelerate the search process by reusing proxy search outcomes for similar external neighbors. Furthermore, we provide a theoretical analysis and show that using a proxy node has a similar influence on training when it is sufficiently similar to the external one. Extensive evaluations show thatGproxysignificantly reduces communication cost while maintaining model performance compared to strong baselines. Junyang Wang 0004, Lan Zhang 0002, Mu Yuan, Yihang Cheng 0002, Yunhao Yao, Zhonghao Hu |
IEEE Trans. Mob. Comput. | 4 |
| 2026 | PrivGuardInfer: Channel-Level End-Edge Collaborative Inference Strategy Protecting Original Inputs and Sensitive AttributesabstractEnd-edge collaborative inference improves computational efficiency by dividing a deep neural network into two parts, executed across the end device and the edge node in parallel. However, adversaries like malicious edge nodes can exploit transmitted data to reconstruct original inputs or infer sensitive attributes. Existing collaborative inference strategies upload the majority of input features to the edge node, significantly increasing the risk of privacy leakage, even without input reconstruction. Therefore, we propose PrivGuardInfer, a channel-level DNN end-edge collaborative inference strategy that optimizes intra-layer partition to simultaneously protect original inputs and sensitive attributes while ensuring latency constraints, supported by three key designs. First, the privacy measurements oriented both layer depth and channel count, jointly quantify the difficulty of reconstructing original inputs using varying numbers of feature maps across different layers. After assessing each channel's contribution, the information offset further measures the difficulty of inferring sensitive attributes. Finally, PrivGuardInfer models the privacy-optimal intra-layer partition under latency constraints as a grouped knapsack problem, mapping attack difficulty to item values and inference latency to item weights. Experimental results reveal that PrivGuardInfer achieves an average improvement of 80.54% in defending against model inversion attacks and 63.34% against attribute inference attacks compared to existing end-edge partition strategies. Moreover, it outperforms current privacy protection methods by an average of 69.37% and 49.75% in mitigating these two types of attacks. Yunhao Yao, Puhan Luo, Yihang Cheng 0002, Jiahui Hou, Xiang-Yang Li 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | FedMix: Boosting with Data Mixture for Vertical Federated LearningabstractThe need to safeguard data privacy and adhere to regulations such as GDPR creates data silos and has prompted the emergence and widespread adoption of techniques for distributed databases. To effectively explore the value of data across multiple organizations, techniques for data management, data analysis and data functionality from distributed databases have been proposed. Recently, Vertical Federated Learning (VFL) has become a solution with growing interests, which enables collaborative model training when data features are partitioned into multiple parts and are held by different parties. However, typical VFL methods heavily rely on private set intersection (PSI) to align data before training and only utilize aligned data for training. In this work, we provide a theoretical analysis to show that unaligned data actually contains valuable and rich features, and a thoughtful design that harnesses the potential of unaligned samples to significantly improve the performance of VFL models. Regrettably, many existing methods simply discard unaligned data, resulting in an irrecoverable loss of performance. To address this data sacrifice problem, we introduce the concept of data mixture, which enables the utilization of both aligned and unaligned data during training. Building upon the data mixture idea, we present FedMix, the first on-the-fly and distribution-agnostic framework designed to boost the performance of VFL models by leveraging unaligned data. A data seasoning approach is also designed to utilize auxiliary data lacking label information. Evaluations on diverse datasets under different settings demonstrate the effectiveness of the proposed FedMix compared with various SOTA approaches. FedMix achieves up to 15% model performance improvement and 30.5 hours time cost reduction. Yihang Cheng 0002, Lan Zhang 0002, Junyang Wang 0004, Xiaokai Chu, Dongbo Huang, Lan Xu 0001 |
ICDE | 1 |
| 2024 | GraphProxy: Communication-Efficient Federated Graph Learning with Adaptive ProxyabstractFederated graph learning (FGL) enables multiple participants with distributed but connected graph data to collaboratively train a model in a privacy-preserving way. However, the high communication cost hinder the adoption of FGL in many resource-limited or delay-sensitive applications. In this work, we focus on reducing the communication cost incurred by the transmission of neighborhood information in FGL. We propose to search for local proxies that can play a substitute role as the external neighbors, and develop a novel federated graph learning framework named GraphProxy. GraphProxy utilizes representation similarity and class correlation to select local proxies for external neighbors. And we propose to dynamically adjust the proxy strategy according to the changing representation of nodes during the iterative training process. We also perform a theoretical analysis and show that using a proxy node has a similar influence on training when it is sufficiently similar to the external one. Extensive evaluations show the effectiveness of our design, e.g., GraphProxy can achieve 8× communication efficiency with only 0.14% performance degradation. Junyang Wang 0004, Lan Zhang 0002, Mu Yuan, Yihang Cheng 0002 |
INFOCOM | 5 |
| 2024 | SecoInfer: Secure DNN End-Edge Collaborative Inference Framework Optimizing Privacy and LatencyabstractEnd-edge collaborative inference enhances computational efficiency by segmenting a deep neural network (DNN) model into two parts, executed across the end device and the edge node. However, existing collaborative inference strategies often involve transmitting original inputs from the end device to the edge node, resulting in significant risks of user detail leakage without requiring input reconstruction. Therefore, in this work, we present SecoInfer, a secure layer-level DNN end-edge collaborative inference framework. SecoInfer achieves joint optimization of data privacy and inference latency for DNN partition solutions that meet latency constraints, supported by three key designs. First, the privacy-aware DNN layer projection measurement quantifies the difficulty adversaries encounter in reconstructing the original input from the intermediate output of each layer. Then, the latency-privacy integrated structure modeling enables the direct calculation of the privacy measurement and inference latency for each partition solution from a list element or a directed acyclic graph (DAG) cut. Finally, the two-stage latency constraint adjustment scheme narrows down the search space of feasible partition solutions at the block level and fine-tunes the final one to meet the latency constraint based on layer depth. We prototype SecoInfer, utilizing a Raspberry Pi 4B as the end device and a server with an NVIDIA GeForce RTX 3060 GPU as the edge node. Experimental results demonstrate that under latency constraints of 20 ms, 33 ms, and 40 ms, SecoInfer reduces adversarial data reconstruction by 9.84%, 19.26%, and 25.18%, respectively, without any loss of task model accuracy. SecoInfer also enhances efficiency, reducing the time needed to determine optimal end-edge partition solutions on a Raspberry Pi 4B by 18.04%. Yunhao Yao, Jiahui Hou, Yihang Cheng 0002, Mu Yuan, Puhan Luo, Xiang-Yang Li 0001 |
ACM Trans. Sens. Networks | 4 |
| 2023 | TVFL: Tunable Vertical Federated Learning towards Communication-Efficient Model Serving
Lan Zhang 0002, Yihang Cheng 0002, Shaoang Li, Dongbo Huang, Xu Lan |
INFOCOM | 3 |
| 2023 | GFL: Federated Learning on Non-IID Data via Privacy-Preserving Synthetic DataabstractFederated learning (FL) enables large amounts of participants to construct a global learning model, while storing training data privately at local client devices. A fundamental issue in FL systems is the susceptibility to the highly skewed distributed data. A series of methods have been proposed to mitigate the Non-IID problem by limiting the distances between local models and the global model, but they cannot address the root cause of skewed data distribution eventually. Some methods share extra samples from the server to clients, which requires comprehensive data collection by the server and may raise potential privacy risks. In this work, we propose an efficient and adaptive framework, named Generative Federated Learning (GFL), to solve the skewed data problem in FL systems in a privacy-friendly way. We introduce Generative Adversarial Networks (GAN) into FL to generate synthetic data, which can be used by the server to balance data distributions. To keep the distribution and membership of clients' data private, the synthetic samples are generated with random distributions and protected by a differential privacy mechanism. The results show that GFL significantly outperforms existing approaches in terms of achieving more accurate global models (e.g., 17% - 50% higher accuracy) as well as building global models with faster convergence speed without increasing much computation or communication costs. Yihang Cheng 0002, Lan Zhang 0002, Anran Li 0001 |
PERCOM | 1 |