VLDB 2026 Research / reviewers in the wild / expert
Junyang Wang 0004
dblp:254/9084-4
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0006-0716-136XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TopFGL: A Topology-Aware and Distributionagnostic Federated Learning Framework Tackling Topological Heterogeneity on Graph Data
Junyang Wang 0004, Lan Zhang 0002, Yihang Cheng 0002, Mu Yuan, Tian Wang 0001, Zhihui Fu |
ICDE | 1 |
| 2026 | Leveraging Confidence Consistency for Poisoned Sample Detection in Vertical Federated LearningabstractVertical federated learning (VFL) enables multiple parties with non-overlapping features to collaboratively train a model without sharing raw data. However, its split-learning architecture makes VFL particularly vulnerable to targeted backdoor attacks, while also limiting access to raw features and complete model parameters. These constraints render defenses developed for horizontal federated learning inapplicable. Moreover, existing VFL defenses neglect the defender’s privileged role as the label holder, leaving them vulnerable to adaptive backdoor strategies. In this paper, we propose CONFIRM, a post-training defense framework that leverages this privileged position to detect poisoned samples. Our approach is driven by a key insight: poisoned samples exhibit abnormally stable prediction confidence under model perturbations, an intrinsic property of backdoor attacks. By actively perturbing the top model and measuring the consistency of prediction confidence, CONFIRM detects poisoned samples based on their perturbation robustness. Extensive experiments across multiple benchmarks show that CONFIRM consistently outperforms state-of-the-art methods, achieving AUROC scores exceeding 99.9% and up to 21.2% higher F1-scores. Furthermore, CONFIRM remains effective against adaptive attacks. We believe that our findings highlight a promising direction for backdoor defense in VFL, and pave the way for more secure VFL systems. Lan Zhang 0002, Quanchao Liu, Yetian He, Junyang Wang 0004, Peng Ran |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2026 | Gproxy: Communication-Efficient Federated Graph Learning With Efficient Adaptive ProxyingabstractFederated graph learning (FGL) enables multiple participants with distributed but connected graph data to collaboratively train a model in a privacy-preserving way. However, the high communication cost hinders the adoption of FGL in many resource-limited or delay-sensitive applications. In this work, we focus on reducing the communication cost incurred by the transmission of neighborhood information in FGL. We propose to search for local proxies that can play a substitute role as the external neighbors and develop a novel federated graph learning framework namedGproxy.Gproxyutilizes representation similarity and class correlation to select local proxies for external neighbors. Additionally, we propose to dynamically adjust the proxy strategy according to the changing representation of nodes during the iterative training process. We also design a proxy cache to accelerate the search process by reusing proxy search outcomes for similar external neighbors. Furthermore, we provide a theoretical analysis and show that using a proxy node has a similar influence on training when it is sufficiently similar to the external one. Extensive evaluations show thatGproxysignificantly reduces communication cost while maintaining model performance compared to strong baselines. Junyang Wang 0004, Lan Zhang 0002, Mu Yuan, Yihang Cheng 0002, Yunhao Yao, Zhonghao Hu |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | FedMix: Boosting with Data Mixture for Vertical Federated LearningabstractThe need to safeguard data privacy and adhere to regulations such as GDPR creates data silos and has prompted the emergence and widespread adoption of techniques for distributed databases. To effectively explore the value of data across multiple organizations, techniques for data management, data analysis and data functionality from distributed databases have been proposed. Recently, Vertical Federated Learning (VFL) has become a solution with growing interests, which enables collaborative model training when data features are partitioned into multiple parts and are held by different parties. However, typical VFL methods heavily rely on private set intersection (PSI) to align data before training and only utilize aligned data for training. In this work, we provide a theoretical analysis to show that unaligned data actually contains valuable and rich features, and a thoughtful design that harnesses the potential of unaligned samples to significantly improve the performance of VFL models. Regrettably, many existing methods simply discard unaligned data, resulting in an irrecoverable loss of performance. To address this data sacrifice problem, we introduce the concept of data mixture, which enables the utilization of both aligned and unaligned data during training. Building upon the data mixture idea, we present FedMix, the first on-the-fly and distribution-agnostic framework designed to boost the performance of VFL models by leveraging unaligned data. A data seasoning approach is also designed to utilize auxiliary data lacking label information. Evaluations on diverse datasets under different settings demonstrate the effectiveness of the proposed FedMix compared with various SOTA approaches. FedMix achieves up to 15% model performance improvement and 30.5 hours time cost reduction. Yihang Cheng 0002, Lan Zhang 0002, Junyang Wang 0004, Xiaokai Chu, Dongbo Huang, Lan Xu 0001 |
ICDE | 3 |
| 2024 | GraphProxy: Communication-Efficient Federated Graph Learning with Adaptive ProxyabstractFederated graph learning (FGL) enables multiple participants with distributed but connected graph data to collaboratively train a model in a privacy-preserving way. However, the high communication cost hinder the adoption of FGL in many resource-limited or delay-sensitive applications. In this work, we focus on reducing the communication cost incurred by the transmission of neighborhood information in FGL. We propose to search for local proxies that can play a substitute role as the external neighbors, and develop a novel federated graph learning framework named GraphProxy. GraphProxy utilizes representation similarity and class correlation to select local proxies for external neighbors. And we propose to dynamically adjust the proxy strategy according to the changing representation of nodes during the iterative training process. We also perform a theoretical analysis and show that using a proxy node has a similar influence on training when it is sufficiently similar to the external one. Extensive evaluations show the effectiveness of our design, e.g., GraphProxy can achieve 8× communication efficiency with only 0.14% performance degradation. Junyang Wang 0004, Lan Zhang 0002, Mu Yuan, Yihang Cheng 0002 |
INFOCOM | 1 |
| 2023 | ENLD: Efficient Noisy Label Detection for Incremental Datasets in Data LakeabstractDue to the difficulty of obtaining high-quality data in real-world scenarios, datasets inevitably contain noisy labeled data, leading to inefficient data usage and poor model performance. Thus, noisy label detection is an important research topic. Previous efforts mainly focus on noisy label detection on specific datasets that have been collected. Some works select clean samples based on relations between representations during the training process; some works utilize confidence outputs of a pre-trained model for noisy label detection. However, how to perform efficient and fine-grained noisy label detection on constantly arriving datasets in a data lake with a large amount of inventory data has not been explored. The rapidly growing volume and changing distribution of data make conventional methods either incur large computation overhead due to repeated training or become increasingly ineffective on newly arriving data. To address these challenges, in this work, we propose a novel approach ENLD to perform efficient and accurate noisy label detection on incremental datasets. Our extensive experiments demonstrate that ENLD outperforms the next best method in both efficiency and accuracy, which achieves 3.65 ×-4.97× detection speedup and higher average f1 scores with various noise rate settings. Xuanke You, Lan Zhang 0002, Junyang Wang 0004, Zhimin Bao, Shuaishuai Dong |
ICDE | 3 |
| 2021 | Visual SLAM for robot navigation in healthcare facility
Baofu Fang, Gaofei Mei, Xiaohui Yuan 0001, Zaijun Wang, Junyang Wang 0004 |
Pattern Recognit. | 6 |