VLDB 2026 Research / reviewers in the wild / expert
Shenbao Yu
dblp:198/6469
· DBLP profile ↗
21ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0002-6824-4293ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 12 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BiO-HMC: Dynamic Human-Machine Collaboration for Consensus Decision-Making via Bilevel OptimizationabstractConsensus decision-making uses crowd responses (usually from non-experts) to questions to reach a consensus answer based on human-machine collaboration. The crucial point is dynamic, which should not only enable rapid self-iteration toward the correct answer through crowd workers' responses but also adaptively suggest the next most valuable question(s) to accelerate the integration of the answer. However, existing methods reach consensus using either offline data or fixed question search structures, thereby largely sidestepping this dynamic nature. In response, we propose a bilevel optimization-based human-machine collaboration (BiO-HMC), which explores an inner & outer-level optimization to enable effective answer integration and efficient question selection. The resulting optimization problem is intractable because there is no closed-form expression in the inner-level optimization. We employ a gradient-based method and guarantee the method's theoretical convergence. Experimental results on synthetic and real-world datasets demonstrate the effectiveness and efficiency of the BiO-HMC model, i.e., achieving the highest confidence in the correct answer with the lowest labor cost. Yinghui Pan, Shuaijie Zhao, Shenbao Yu, Zongyang Liu, Yifeng Zeng, Han Liu 0002, Mingwei Lin |
AAAI | 3 |
| 2026 | GBCG: Granular Ball and Counterfactual Guided Profile Injection Attack in Recommender Systems
Yunmeng Zhao, Yuran He, Shenbao Yu, Ruihong Huang, Jun Shen 0001, Jiayin Lin |
DASFAA (1) | 3 |
| 2026 | A time-aware enhanced contrastive learning framework for mitigating popularity bias in sequential recommender systems
Zhewei Ding, Shenbao Yu, Jun Shen 0001, Jiayin Lin |
Expert Syst. Appl. | 2 |
| 2026 | A Lightweight and Robust Low-Rank Tensor Embedding-Integrated VAE for Network Anomaly Detection
Mingwei Lin, Shenbao Yu, Riqing Chen, Xin Luo 0001 |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2025 | Dual-Generator Augmentation for Imbalanced Fraud Detection on Heterogeneous Insurance GraphsabstractFraud detection holds paramount significance across numerous fields, especially in health insurance. Boosted by medical users' rich information, which is featured with multimodality and often takes the form of heterogeneous graphs, the fraud detection can be intuitively formulated as a graph-based classification task, but it holds a shortcoming: the class imbalance problem (i.e., the number of fraudsters is far less than the normal ones). While efforts in graph-based augmentation have been explored to gain prolific information to improve the class balance, most of these methods scale up only to homogeneous graphs, which do not cater to heterogeneous ones due to the inevitable information loss. Hence, the heterogeneous graph-based augmentation for fraud detection is still largely under-explored. In response, we propose a heterogeneous graph node-level augmentation (HGNA) model to alleviate the negative impact of class imbalance. Specifically, we first design an attribute generator to create synthetic users of the minority class and then propose a topology generator to produce reasonable topological connections. Experiments on public and private real-world datasets demonstrate the efficacy of the HGNA model, suggesting that it can serve as an augmentation proxy to benefit downstream classifiers for fraud detection tasks. Zijie Bian, Kaibiao Lin, Shenbao Yu |
BIBM | 4 |
| 2025 | Rethinking Privacy Protection for Recommender System in a Collaborative WayabstractRecommender systems play a crucial role in personalizing user experiences by analyzing vast amounts of user interaction records to suggest relevant items. However, the use of sensitive user information raises significant privacy concerns, particularly in light of regulations such as the General Data Protection Regulation. In response to these challenges, this paper introduces Rapid Collaborative Machine Unlearning (RaCoMU), a novel framework designed to efficiently handle user data deletion requests while preserving recommendation quality. RaCoMU employs a two-step data partitioning strategy that consolidates user data within individual subsets, enhancing collaborative information utilization. Additionally, a collaboration-based attention aggregation method is proposed, which weights contributions of the submodels based on user similarities. Our extensive experiments on real-world datasets across three recommendation models demonstrate that the proposed RaCoMU achieves superior unlearning efficiency and significantly outperforms existing frameworks in model effectiveness. The findings underscore the potential of RaCoMU to balance user privacy and recommendation performance effectively. Yongpei Zhang, Mingwei Lin, Jun Shen 0001, Shenbao Yu, Jiayin Lin |
CSCWD | 5 |
| 2025 | Item Popularity Attention for Mitigating Popularity Bias in Sequential Recommendation
Zhewei Ding, Runxiong Liu, Mingwei Lin, Shenbao Yu, Jiayin Lin |
ICONIP (1) | 4 |
| 2025 | EPCTS: Enhanced Prompt-Aware Cross-Prompt Essay Trait Scoring
Jiangsong Xu, Mingwei Lin, Jiayin Lin, Shenbao Yu, Liang Zhao 0004, Jun Shen 0001 |
Neurocomputing | 5 |
| 2025 | An Incremental Nonlinear Co-Latent Factor Analysis Model for Large-Scale Student Performance PredictionabstractPredicting student performance (PSP) is critical to intelligent tutoring systems in online education services. Accurate predictions enable data-driven decision making and facilitate the implementation of timely educational interventions. While previous approaches, such as cognitive diagnosis models and data mining techniques, have demonstrated effectiveness with small-scale and static datasets, they face significant challenges in large-scale online learning environments. In these settings, vast volumes of learning responses are continuously generated as students engage with exercises. Those responses are characterized by high dimensionality and incompleteness (HDI) and often manifest as streaming data, which limits the applicability of most existing prediction methods, therefore posing a new challenge to traditional PSP tasks. To remedy the void of PSP in the HDI and incremental scenario, we propose an incremental nonlinear co-latent factor analysis (IN-CoLFA) model, which enhances latent factor analysis – an element-wise learning framework – by incorporating a co-factorization technique and integrating a neural network-inspired structure. To facilitate incremental learning, we develop momentum-accelerated stochastic gradient-based algorithms, which enable the model to perform offline training on historical data and continuously refine its predictions as new student performance becomes available. Experiments on several real-world datasets demonstrate the efficacy and efficiency of our approach in both stationary and incremental PSP tasks under HDI conditions. Shenbao Yu, Mingwei Lin, Xiuqin Xu, Jiayin Lin, Zeshui Xu |
IEEE Trans. Serv. Comput. | 2 |
| 2024 | Causal-Driven Skill Prerequisite Structure DiscoveryabstractKnowing a prerequisite structure among skills in a subject domain effectively enables several educational applications, including intelligent tutoring systems and curriculum planning. Traditionally, educators or domain experts use intuition to determine the skills' prerequisite relationships, which is time-consuming and prone to fall into the trap of blind spots. In this paper, we focus on inferring the prerequisite structure given access to students' performance on exercises in a subject. Nevertheless, it is challenging since students' mastery of skills can not be directly observed, but can only be estimated, i.e., its latency in nature. To tackle this problem, we propose a causal-driven skill prerequisite structure discovery (CSPS) method in a two-stage learning framework. In the first stage, we learn the skills' correlation relationships presented in the covariance matrix from the student performance data while, through the predicted covariance matrix in the second stage, we consider a heuristic method based on conditional independence tests and standardized partial variance to discover the prerequisite structure. We demonstrate the performance of the new approach with both simulated and real-world data. The experimental results show the effectiveness of the proposed model for identifying the skills' prerequisite structure. Shenbao Yu, Yifeng Zeng, Fan Yang 0010, Yinghui Pan |
AAAI | 1 |
| 2024 | An Autoencoder-Like Nonnegative Matrix Co-Factorization for Improved Student Cognitive ModelingabstractStudent cognitive modeling (SCM) is a fundamental task in intelligent education, with applications ranging from personalized learning to educational resource allocation. By exploiting students' response logs, SCM aims to predict their exercise performance as well as estimate knowledge proficiency in a subject. Data mining approaches such as matrix factorization can obtain high accuracy in predicting student performance on exercises, but the knowledge proficiency is unknown or poorly estimated. The situation is further exacerbated if only sparse interactions exist between exercises and students (or knowledge concepts). To solve this dilemma, we root monotonicity (a fundamental psychometric theory on educational assessments) in a co-factorization framework and present an autoencoder-like nonnegative matrix co-factorization (AE-NMCF), which improves the accuracy of estimating the student's knowledge proficiency via an encoder-decoder learning pipeline. The resulting estimation problem is nonconvex with nonnegative constraints. We introduce a projected gradient method based on block coordinate descent with Lipschitz constants and guarantee the method's theoretical convergence. Experiments on several real-world data sets demonstrate the efficacy of our approach in terms of both performance prediction accuracy and knowledge estimation ability, when compared with existing student cognitive models. Shenbao Yu, Yinghui Pan, Yifeng Zeng, Prashant Doshi, Guoquan Liu, Kim-Leng Poh, Mingwei Lin |
NeurIPS | 1 |
| 2024 | CA-PDBPR: category-aware privacy preserving POI recommendation using decentralized Bayesian personalized ranking
Qinyun Gao, Shenbao Yu, Bilian Chen, Langcai Cao |
Appl. Intell. | 2 |
| 2024 | Revisiting the loss functions in sequential recommendation
Fangyu Li 0003, Shenbao Yu, Fan Yang 0010 |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | A hybrid similarity model for mitigating the cold-start problem of collaborative filtering in sparse data
Jiewen Guan, Bilian Chen, Shenbao Yu |
Expert Syst. Appl. | 3 |
| 2024 | A substructure transfer reinforcement learning method based on metric learning
Peihua Chai, Bilian Chen, Yifeng Zeng, Shenbao Yu |
Neurocomputing | 4 |
| 2024 | Overlapping community detection using expansion with contraction
Bilian Chen, Shenbao Yu, Langcai Cao |
Neurocomputing | 3 |
| 2024 | Improving collaborative filtering with SNE-GCN: a second-order neighbor enhanced graph convolutional network
Tianyang Yan, Langcai Cao, Peihua Chai, Shenbao Yu |
Multim. Syst. | 4 |
| 2024 | SNMCF: A Scalable Non-Negative Matrix Co-Factorization for Student Cognitive ModelingabstractStudent cognitive modeling plays an important role in the rapid development of educational data mining research. It aims to discover students' proficiency in knowledge concepts as well as to predict students' performance in conducting exercises. Studies in the past few years have been mainly centered around two types of techniques: cognitive diagnosis models and data mining approaches. Cognitive diagnosis models focus on students' cognitive states and assess their knowledge concept proficiency through handcrafted features. The subjective features may trigger cascading errors in the students' performance prediction. On the other hand, data mining techniques, e.g., matrix factorization methods, achieve high prediction accuracy by directly modeling the students' exercising process. It lacks measuring the students' knowledge concept proficiency. To address the dilemma of the aforementioned methods, in this paper, we propose a scalable non-negative matrix co-factorization (SNMCF) model by jointly modeling the students' knowledge states and their exercising process. SNMCF can achieve high accuracy in predicting students' exercise performance while modeling their states of knowledge concepts in a given domain. We conduct extensive experiments on several real-world datasets, including large sparse ones, and demonstrate the effectiveness of our new approach in terms of prediction accuracy, cognitive diagnostic ability, and scalability Shenbao Yu, Yifeng Zeng, Yinghui Pan, Fan Yang 0010 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Discovering a cohesive football team through players' attributed collaboration networksabstractAbstract The process of team composition in multiplayer sports such as football has been a main area of interest within the field of the science of teamwork, which is important for improving competition results and game experience. Recent algorithms for the football team composition problem take into account the skill proficiency of players but not the interactions between players that contribute to winning the championship. To automate the composition of a cohesive team, we consider the internal collaborations among football players. Specifically, we propose a Team Composition based on the Football Players’ Attributed Collaboration Network (TC-FPACN) model, aiming to identify a cohesive football team by maximizing football players’ capabilities and their collaborations via three network metrics, namely, network ability, network density and network heterogeneity&homogeneity. Solving the optimization problem is NP-hard; we develop an approximation method based on greedy algorithms and then improve the method through pruning strategies given a budget limit. We conduct experiments on two popular football simulation platforms. The experimental results show that our proposed approach can form effective teams that dominate others in the majority of simulated competitions. Shenbao Yu, Yifeng Zeng, Yinghui Pan, Bilian Chen |
Appl. Intell. | 1 |
| 2023 | Generalized temporal similarity-based nonnegative tensor decomposition for modeling transition matrix of dynamic collaborative filtering
Shenbao Yu, Zhehao Zhou, Bilian Chen, Langcai Cao |
Inf. Sci. | 1 |
| 2017 | Using function approximation for personalized point-of-interest recommendation
Bilian Chen, Shenbao Yu, Jing Tang 0001, Mengda He, Yifeng Zeng |
Expert Syst. Appl. | 2 |