VLDB 2026 Research / reviewers in the wild / expert
Yuhang Yao 0003
dblp:203/0159-3
· DBLP profile ↗
11ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0002-7045-0002ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Computer networks · 4Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Federated Large Language Models: Current Progress and Future Directions
Yuhang Yao 0003, Junda Wu, Chengkai Huang, Yu Xia 0007, Tong Yu 0001, Ruiyi Zhang 0002, Sungchul Kim, Ryan Rossi, Ang Li 0005, Lina Yao 0001, Julian J. McAuley, Yiran Chen 0001, Carlee Joe-Wong |
PAKDD (4) | 1 |
| 2024 | FedSecurity: A Benchmark for Attacks and Defenses in Federated Learning and Federated LLMsabstractThis paper introduces FedSecurity, an end-to-end benchmark that serves as a supplementary component of the FedML library for simulating adversarial attacks and corresponding defense mechanisms in Federated Learning (FL). FedSecurity eliminates the need for implementing the fundamental FL procedures, e.g., FL training and data loading, from scratch, thus enables users to focus on developing their own attack and defense strategies. It contains two key components, including FedAttacker that conducts a variety of attacks during FL training, and FedDefender that implements defensive mechanisms to counteract these attacks. FedSecurity has the following features: i) It offers extensive customization options to accommodate a broad range of machine learning models (e.g., Logistic Regression, ResNet, and GAN) and FL optimizers (e.g., FedAVG, FedOPT, and FedNOVA); ii) it enables exploring the effectiveness of attacks and defenses across different datasets and models; and iii) it supports flexible configuration and customization through a configuration file and some APIs. We further demonstrate FedSecurity's utility and adaptability through federated training of Large Language Models (LLMs) to showcase its potential on a wide range of complex applications. Baturalp Buyukates, Zijian Hu 0001, Weizhao Jin, Lichao Sun 0001, Chulin Xie, Yuhang Yao 0003, Kai Zhang 0039, Qifan Zhang 0002, Carlee Joe-Wong, Amir Salman Avestimehr, Chaoyang He 0001 |
KDD | 10 |
| 2023 | Wyze Rule: Federated Rule Dataset for Rule Recommendation BenchmarkingabstractIn the rapidly evolving landscape of smart home automation, the potential of IoT devices is vast. In this realm, rules are the main tool utilized for this automation, which are predefined conditions or triggers that establish connections between devices, enabling seamless automation of specific processes. However, one significant challenge researchers face is the lack of comprehensive datasets to explore and advance the field of smart home rule recommendations. These datasets are essential for developing and evaluating intelligent algorithms that can effectively recommend rules for automating processes while preserving the privacy of the users, as it involves personal information about users' daily lives. To bridge this gap, we present the Wyze Rule Dataset, a large-scale dataset designed specifically for smart home rule recommendation research. Wyze Rule encompasses over 1 million rules gathered from a diverse user base of 300,000 individuals from Wyze Labs, offering an extensive and varied collection of real-world data. With a focus on federated learning, our dataset is tailored to address the unique challenges of a cross-device federated learning setting in the recommendation domain, featuring a large-scale number of clients with widely heterogeneous data. To establish a benchmark for comparison and evaluation, we have meticulously implemented multiple baselines in both centralized and federated settings. Researchers can leverage these baselines to gauge the performance and effectiveness of their rule recommendation systems, driving advancements in the domain. The Wyze Rule Dataset is publicly accessible through HuggingFace's dataset API. Mohammad Mahdi Kamani, Yuhang Yao 0003, Hanjia Lyu, Zhongwei Cheng, Lin Chen 0021, Liangju Li, Carlee Joe-Wong, Jiebo Luo 0001 |
NeurIPS | 2 |
| 2023 | FedGCN: Convergence-Communication Tradeoffs in Federated Training of Graph Convolutional NetworksabstractMethods for training models on graphs distributed across multiple clients have recently grown in popularity, due to the size of these graphs as well as regulations on keeping data where it is generated. However, the cross-client edges naturally exist among clients. Thus, distributed methods for training a model on a single graph incur either significant communication overhead between clients or a loss of available information to the training. We introduce the Federated Graph Convolutional Network (FedGCN) algorithm, which uses federated learning to train GCN models for semi-supervised node classification with fast convergence and little communication. Compared to prior methods that require extra communication among clients at each training round, FedGCN clients only communicate with the central server in one pre-training step, greatly reducing communication costs and allowing the use of homomorphic encryption to further enhance privacy. We theoretically analyze the tradeoff between FedGCN's convergence rate and communication cost under different data distributions. Experimental results show that our FedGCN algorithm achieves better model accuracy with 51.7\% faster convergence on average and at least 100$\times$ less communication compared to prior work. Yuhang Yao 0003, Weizhao Jin, Srivatsan Ravi, Carlee Joe-Wong |
NeurIPS | 1 |
| 2021 | Interpretable Clustering on Dynamic Graphs with Recurrent Graph Neural NetworksabstractWe study the problem of clustering nodes in a dynamic graph, where the connections between nodes and nodes' cluster memberships may change over time, e.g., due to community migration. We first propose a dynamic stochastic block model that captures these changes, and a simple decay-based clustering algorithm that clusters nodes based on weighted connections between them, where the weight decreases at a fixed rate over time. This decay rate can then be interpreted as signifying the importance of including historical connection information in the clustering. However, the optimal decay rate may differ for clusters with different rates of turnover. We characterize the optimal decay rate for each cluster and propose a clustering method that achieves almost exact recovery of the true clusters. We then demonstrate the efficacy of our clustering algorithm with optimized decay rates on simulated graph data. Recurrent neural networks (RNNs), a popular algorithm for sequence learning, use a similar decay-based method, and we use this insight to propose two new RNN-GCN (graph convolutional network) architectures for semi-supervised graph clustering. We finally demonstrate that the proposed architectures perform well on real data compared to state-of-the-art graph clustering algorithms. Yuhang Yao 0003, Carlee Joe-Wong |
AAAI | 1 |
| 2021 | GCN-SE: Attention as Explainability for Node Classification in Dynamic GraphsabstractGraph Convolutional Networks (GCNs) are a popular method from graph representation learning that have proved effective for tasks like node classification. Recent variants on traditional GCN models aim to classify nodes in dynamic graphs whose topologies and node attributes change over time, e.g., social networks with dynamic relationships. These works, however, do not fully address the challenge of flexibly assigning different importance to snapshots of the graph at different times, which depending on the graph dynamics may have more or less predictive power on the labels. We address this challenge by proposing a new method, GCN-SE, that attaches a set of learnable attention weights to graph snapshots at different times, inspired by Squeeze and Excitation Net (SE-Net). We show that GCNSE outperforms previously proposed node classification methods on a variety of graph datasets. To verify the effectiveness of the attention weight in determining the importance of different graph snapshots, we adapt perturbation-based methods from the field of explainable machine learning to graphical settings and evaluate the correlation between the attention weights learned by GCN-SE and the importance of different snapshots over time. Yucai Fan, Yuhang Yao 0003, Carlee Joe-Wong |
ICDM | 2 |
| 2019 | APRP: An Anonymous Propagation Method in Bitcoin Network
Yuhang Yao 0003, Luoyi Fu, Xinbing Wang |
AAAI | 1 |
| 2019 | Modeling, Analysis and Validation of Evolving Networks With Hybrid InteractionsabstractIn many real-world networks, entities of different types usually form an evolving network with hybrid interactions. However, how to theoretically model such networks, along with quantitive characterizations, remains unexplored. Motivated by this, we develop a novel evolving model, which, as validated by our empirical results, can well capture some basic properties such as power-law degree distribution, densification, shrinking diameter, and community structure embodied in most real datasets. Particularly, two types of results are presented in this paper. First, our proposed model, namely, evolving K-Graph, consists of K-node sets representing K different types of entities. The hybrid interactions among entities, based on whether they belong to the same type, are classified into inter-type and intra-type ones that are, respectively, characterized by two joint graphs evolving over time. Following our newly proposed mechanism called interactive-evolution, potential connections can be established among nodes with common features and further form a positive feedback. The superiorities of our model are three folded: good capture of realistic networks, mathematical tractability and efficient implementation. Second, by analytical derivations, along with empirical validation on real datasets, we disclose two aspects of network properties: basic ones as power-law degree distribution, densification, shrinking diameter and community structure, as well as a distinctive one, that is, positive correlation observed in real networks, implying that a hub in one inter-type relationship network also has many neighbors in another one. An additional interesting finding is that through further comparison of models with or without interactive-evolution, the former one leads to an even earlier occurrence of network connectivity. Jiaqi Liu 0002, Luoyi Fu, Yuhang Yao 0003, Xinzhe Fu, Xinbing Wang, Guihai Chen |
IEEE/ACM Trans. Netw. | 3 |
| 2018 | GLP: A Novel Framework for Group-Level Location Promotion in Geo-Social NetworksabstractLocation-aware viral marketing is crucial in modern commercial applications for attracting customers to certain points of interests. Prior works are mainly based on formulating it into a location-aware influence maximization problem in Geo-social Networks (GSNs), where $K$ initial seed individuals are selected in hope of maximizing the number of final influenced users. In this paper, we present the first look into the group-level location promotion, which can potentially enhance its performance, with the phenomenon that users belonging to the same geo-community share similar moving preferences. We propose GLP, a new and novel framework of group-level location promotion by virtue of geo-communities, each of which is treated as a group in GSNs. Aiming to attract more users to designated locations, GLP firstly carries out user grouping through an iterative learning approach based on information extraction from massive check-ins records. The advantage of GLP is three-folded: i) by aggregating movements of group members, GLP significantly avoids the sparsity and sporadicity of individual check-ins, and thus obtains more reliable mobility models; ii) by generalizing a new group-level social graph, GLP can exponentially reduce the computational complexity of seed nodes selection that is algorithmically executed by a greedy algorithm; iii) in comparison with prior individual-level cases, GLP is theoretically demonstrated to drastically increase influence spread under the same given budget. Extensive experiments on real datasets demonstrate that the GLP outperforms four baselines, with notably up to 10 times larger influence spread and 100 times faster seed selection over two individual-level cases, meanwhile verifying the impact of group numbers in final influence spread. Luoyi Fu, Yuhang Yao 0003, Xinzhe Fu, Xinbing Wang, Guihai Chen |
IEEE/ACM Trans. Netw. | 3 |
| 2017 | Evolving K-Graph: Modeling Hybrid Interactions in NetworksabstractIn many realistic networks, entities of different types usually form an evolving network with hybrid interactions. However, how to mathematically model such networks remains unexplored. Motivated by this, we develop a novel evolving model, which, as validated by our empirical results, can well capture some basic features such as power-law distribution, densification and shrinking diameter. Particularly, in our proposed model, named Evolving K-Graph, the hybrid interactions among entities are classified into inter-type and intra-type connections that are respectively characterized by two joint graphs evolving over time. By empirical validation, we disclose two new network properties: a positive correlation of any two layers of the network, and an earlier occurrence of network connectivity resulted by our model. Jiaqi Liu 0002, Yuhang Yao 0003, Xinzhe Fu, Luoyi Fu, Xiao-Yang Liu, Xinbing Wang |
MobiHoc | 2 |
| 2017 | Core Percolation in Coupled NetworksabstractCore percolation is crucial in network controllability and robustness. Prior works are mainly based on single, non-interacting network where core nodes are obtained by a classic Greedy Leaf Removal (GLR) procedure that takes of leaf nodes along with their neighbors iteratively. We take a first look into core percolation in coupled networks with two fully-interdependent networks. To obtain core nodes in both networks, we propose a new algorithm, called Alternating GLR procedure, that recursively switches between networks in carrying out GLR for node removal. We prove that the proposed algorithm can guarantee the uniqueness in the sense that the final remaining nodes and edges in either of the coupled networks remain the same, and then present analytical solutions for the fraction of core nodes that will be ultimately left when the algorithm terminates. Our simulation demonstrates that the presence of core exhibits a jump at the critical point as a first order transition in coupled networks. Jiayu Pan, Yuhang Yao 0003, Luoyi Fu, Xinbing Wang |
MobiHoc | 2 |