EDBT 2026 Demo / reviewers in the wild / expert
Gang Cao 0003
dblp:49/982-3
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2024
0009-0005-1142-0128ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | NC-ALG: Graph-Based Active Learning Under Noisy CrowdabstractGraph Neural Networks (GNNs) have achieved great success in various data mining tasks but they heavily rely on a large number of annotated nodes, requiring considerable human efforts. Despite the effectiveness of existing GNN-based Active Learning (AL) methods, they assume that the annotated labels are always correct, which is contradictory to the error-prone labeling process in a practical crowdsourcing environment. Besides, due to this impractical assumption, existing works only focus on optimizing the node selection in AL but neglect optimizing the labeling process. Therefore, we present NC-ALG, the first GNN-based AL framework that optimizes both the node selection and node labeling process under a noisy crowd. For node selection, NC-ALG introduces a new measurement to model influence reliability and an effective influence maximization objective to select nodes. For node labeling, NC-ALG significantly reduces the labeling cost by considering the model-predicted labels and the labels of mirror nodes. To the best of our knowledge, this is the first attempt to consider GNN-based AL under the practical noisy crowd. Empirical studies on public datasets demonstrate that NC-ALG significantly outperforms existing methods in terms labeling efficiency. Notably, it only takes NC-ALG one-third of the labeling budget that the competitive baseline GRAIN needs to achieve an accuracy of 70.7 % on PubMed. Wentao Zhang 0001, Yexin Wang, Zhenbang You, Yang Li 0106, Gang Cao 0003, Zhi Yang 0001, Bin Cui 0001 |
ICDE | 5 |
| 2023 | Hierarchical Interest Modeling of Long-tailed Users for Click-Through Rate PredictionabstractClick-through rate (CTR) prediction, whose purpose is to predict the probability of a user clicking on an item, plays a pivotal role in recommender systems. Capturing users’ accurate preferences from their historical interactions (e.g., clicks) is an essential step for handling this task and has aroused wide concern in both academia and industry. However, most of the previous methods focus on the users with abundant clicks and ill-serve the users who rarely click or purchase items. Though the ratio of these long-tailed users may be small on popular platforms, such as Amazon and Taobao, they are the majority on the newborn e-commerce company like Lazada. To extract the interests of long-tailed users, several works attempt to integrate the side information, such as demographic features. Nevertheless, these features are usually not available and may even lead to privacy concerns. Therefore, how to utilize the noisy and limited clicks becomes the key challenge.In this paper, we propose a novel model called Hierarchical Interest Modeling (HIM). It hierarchically utilizes long-tailed users’ limited behaviors and captures their preferences from both personalized and group-wise perspectives. HIM consists of two main components, including User Behavior Pyramid (UBP) and User Behavior Clustering (UBC). The UBP module utilizes additional negative feedback to reduce the noises in positive feedback, thus obtaining reliable user personalized representations. Then, the UBC module automatically discovers latent user groups with self-supervised reconstruction loss and learns another interest representation for each user in a group-wise aspect. Extensive experiments on both public and industrial datasets verify the superiority of HIM compared with the state-of-the-art baselines. Moreover, HIM has already been deployed on Lazada recommendation scenario and gains 3.38% on CTR prediction on average on the online A/B test. Our codes are available in https://github.com/xiaojin-nj/HIM. Jin Niu, Lifang Deng, Kaigui Bian, Gang Cao 0003, Bin Cui 0001 |
ICDE | 8 |
| 2023 | FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device PlacementabstractWith the increasing data volume, there is a trend of using large-scale pre-trained models to store the knowledge into an enormous number of model parameters. The training of these models is composed of lots of dense algebras, requiring a huge amount of hardware resources. Recently, sparsely-gated Mixture-of-Experts (MoEs) are becoming more popular and have demonstrated impressive pretraining scalability in various downstream tasks. However, such a sparse conditional computation may not be effective as expected in practical systems due to the routing imbalance and fluctuation problems. Generally, MoEs are becoming a new data analytics paradigm in the data life cycle and suffering from unique challenges at scales, complexities, and granularities never before possible. In this paper, we propose a novel DNN training framework, FlexMoE, which systematically and transparently address the inefficiency caused by dynamic dataflow. We first present an empirical analysis on the problems and opportunities of training MoE models, which motivates us to overcome the routing imbalance and fluctuation problems by a dynamic expert management and device placement mechanism. Then we introduce a novel scheduling module over the existing DNN runtime to monitor the data flow, make the scheduling plans, and dynamically adjust the model-to-hardware mapping guided by the real-time data traffic. A simple but efficient heuristic algorithm is exploited to dynamically optimize the device placement during training. We have conducted experiments on both NLP models (e.g., BERT and GPT) and vision models (e.g., Swin). And results show FlexMoE can achieve superior performance compared with existing systems on real-world workloads --- FlexMoE outperforms DeepSpeed by 1.70x on average and up to 2.10x, and outperforms FasterMoE by 1.30x on average and up to 1.45x. Xiaonan Nie, Xupeng Miao, Zilong Wang 0033, Jilong Xue, Lingxiao Ma, Gang Cao 0003, Bin Cui 0001 |
Proc. ACM Manag. Data | 7 |
| 2023 | P2CG: a privacy preserving collaborative graph neural network training framework
Xupeng Miao, Wentao Zhang 0001, Yuezihan Jiang, Fangcheng Fu, Yingxia Shao, Lei Chen 0002, Yangyu Tao, Gang Cao 0003, Bin Cui 0001 |
VLDB J. | 8 |