VLDB 2026 Research / reviewers in the wild / expert
Boyu Li 0003
dblp:25/5732-3
· DBLP profile ↗
19ranked-venue papers
9as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning unified market interdependencies via networked attention for stock price forecastingabstractStock price forecasting is challenging due to market volatility and complex, dynamic inter-stock relationships. Existing graph-based approaches often rely primarily on static relational structures or simple time-aligned correlations, which capture only fixed or short-term relationships and fail to model transient cross-temporal dependencies and heterogeneous information sources. We propose a novel artificial intelligence framework for stock price forecasting that integrates heterogeneous data sources, including price co-movements, corporate linkages derived from Wikipedia, and industry affiliations, into a unified dynamic relational graph that combines both structural and behavioral dependencies. The proposed model, named Learning Unified Market Interdependencies (LUMI), adaptively models evolving inter-stock connections and uncovers latent dependencies beyond sectoral or time-aligned patterns. A dual-path temporal attention mechanism disentangles long-term trends from short-term fluctuations, capturing both periodic behaviors and abrupt market shifts. Extensive experiments on four market datasets demonstrate that the proposed deep learning framework outperforms strong baselines in predictive accuracy while providing interpretable insights into market interdependencies. These findings highlight the potential of artificial intelligence for modeling complex financial systems and improving algorithmic stock price forecasting. Kaveesha Hewage, Boyu Li 0003, Ting Guo 0005, Alexis Stenfors, Peter Mere, Fang Chen 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Cradle: Empowering Foundation Agents towards General Computer ControlabstractDespite their success in specific scenarios, existing foundation agents still struggle to generalize across various virtual scenarios, mainly due to the dramatically different encapsulations of environments with manually designed observation and action spaces. To handle this issue, we propose the General Computer Control (GCC) setting to restrict foundation agents to interact with software through the most unified and standardized interface, i.e., using screenshots as input and keyboard and mouse actions as output. We introduce Cradle, a modular and flexible LMM-powered framework, as a preliminary attempt towards GCC. Enhanced by six key modules, Information Gathering, Self-Reflection, Task Inference, Skill Curation, Action Planning, and Memory, Cradle is able to understand input screenshots and output executable code for low-level keyboard and mouse control after high-level planning and information retrieval, so that Cradle can interact with any software and complete long-horizon complex tasks without relying on any built-in APIs. Experimental results show that Cradle exhibits remarkable generalizability and impressive performance across four previously unexplored commercial video games (Red Dead Redemption 2, Cities:Skylines, Stardew Valley and Dealer’s Life 2), five software applications (Chrome, Outlook, Feishu, Meitu and CapCut), and a comprehensive benchmark, OSWorld. With a unified interface to interact with any software, Cradle greatly extends the reach of foundation agents thus paving the way for generalist agents. Weihao Tan, Wentao Zhang 0007, Xinrun Xu, Haochong Xia, Ziluo Ding, Boyu Li 0003, Junpeng Yue, Jiechuan Jiang, Yewen Li, Ruyi An, Molei Qin, Chuqiao Zong, Longtao Zheng, Xiaoqiang Chai, Yifei Bi, Tianbao Xie, Pengjie Gu, Xiyun Li, Ceyao Zhang, Chaojie Wang 0001, Xinrun Wang, Börje Karlsson 0001, Bo An 0001, Shuicheng Yan, Zongqing Lu 0002 |
ICML | 6 |
| 2025 | Cross-Domain Random Pretraining With Prototypes for Reinforcement LearningabstractUnsupervised cross-domain reinforcement learning (RL) pretraining shows great potential for challenging continuous visual control but poses a big challenge. In this article, we propose cross-domain random pretraining with prototypes (CRPTpro), a novel, efficient, and effective self-supervised cross-domain RL pretraining framework. CRPTpro decouples data sampling from encoder pretraining, proposing decoupled random collection to easily and quickly generate a qualified cross-domain pretraining dataset. Moreover, a novel prototypical self-supervised algorithm is proposed to pretrain an effective visual encoder that is generic across different domains. Without finetuning, the cross-domain encoder can be implemented for challenging downstream tasks defined in different domains, either seen or unseen. Compared with recent advanced methods, CRPTpro achieves better performance on downstream policy learning without extra training on exploration agents for data collection, greatly reducing the burden of pretraining. We conduct extensive experiments across multiple challenging continuous visual-control domains, including balance control, robot locomotion, and manipulation. CRPTpro significantly outperforms the next best Proto-RL(C) on 11/12 cross-domain downstream tasks with only 54.5% wall-clock pretraining time, exhibiting state-of-the-art pretraining performance with greatly improved pretraining efficiency. Xin Liu 0039, Yaran Chen, Haoran Li 0010, Boyu Li 0003, Dongbin Zhao |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2024 | Mitigating Label Bias in Machine Learning: Fairness through Confident LearningabstractDiscrimination can occur when the underlying unbiased labels are overwritten by an agent with potential bias, resulting in biased datasets that unfairly harm specific groups and cause classifiers to inherit these biases. In this paper, we demonstrate that despite only having access to the biased labels, it is possible to eliminate bias by filtering the fairest instances within the framework of confident learning. In the context of confident learning, low self-confidence usually indicates potential label errors; however, this is not always the case. Instances, particularly those from underrepresented groups, might exhibit low confidence scores for reasons other than labeling errors. To address this limitation, our approach employs truncation of the confidence score and extends the confidence interval of the probabilistic threshold. Additionally, we incorporate with co-teaching paradigm for providing a more robust and reliable selection of fair instances and effectively mitigating the adverse effects of biased labels. Through extensive experimentation and evaluation of various datasets, we demonstrate the efficacy of our approach in promoting fairness and reducing the impact of label bias in machine learning models. Yixuan Zhang 0006, Boyu Li 0003, Zenan Ling, Feng Zhou 0011 |
AAAI | 2 |
| 2024 | TransFeat-TPP: An Interpretable Deep Covariate Temporal Point ProcessesabstractThe classical temporal point process (TPP) constructs an intensity function by taking the occurrence times into account. Nevertheless, occurrence time may not be the only relevant factor, other contextual data, termed covariates, may also impact the event evolution. Incorporating such covariates into the model is beneficial, while distinguishing their relevance to the event dynamics is of great practical significance. In this work, we propose a Transformer-based covariate temporal point process (TransFeat-TPP) model to improve the interpretability of deep covariate-TPPs while maintaining powerful expressiveness. TransFeat-TPP can effectively model complex relationships between events and covariates, and provide enhanced interpretability by discerning the importance of various covariates. Experimental results on synthetic and real datasets demonstrate improved prediction accuracy and consistently interpretable feature importance when compared to existing deep covariate-TPPs. Our code is available at https://github.com/waystogetthere/TransFeat.git. Zizhuo Meng, Boyu Li 0003, Xuhui Fan 0001, Zhidong Li, Yang Wang 0002, Fang Chen 0001, Feng Zhou 0011 |
ECAI | 2 |
| 2024 | Neighborhood convolutional graph neural network
Jinsong Chen 0002, Boyu Li 0003, Kun He 0001 |
Knowl. Based Syst. | 2 |
| 2024 | PAMT: A Novel Propagation-Based Approach via Adaptive Similarity Mask for Node ClassificationabstractSemisupervised node classification on attributed networks is a crucial task for network analysis. By decoupling two critical operations in graph convolutional networks (GCNs), namely feature transformation and neighborhood aggregation, recent works of decoupled GCNs could support the information to propagate deeper and achieve advanced performance on node classification. However, they follow the structure-aware propagation strategy of GCNs, making it hard to capture the attribute correlation of nodes and be sensitive to the structure noise described by edges whose two endpoints belong to different categories. To address these issues, we propose a new method called the propagation with adaptive mask then training (PAMT). The key idea is to integrate the attribute similarity mask into the structure-aware propagation process. In this way, PAMT could preserve the attribute correlation of adjacent nodes during the propagation and effectively reduce the influence of structure noise. Moreover, we develop an iterative refinement mechanism to update the similarity mask during the training process to improve the training performance. Extensive experiments on six real-world datasets demonstrate the superior performance and robustness of PAMT over the state-of-the-art baselines. Jinsong Chen 0002, Boyu Li 0003, Qiuting He, Kun He 0001 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2024 | Centroid-Based Multiple Local Community DetectionabstractIn recent years, the research of local community detection has attracted much attention. Most existing local community detection methods aim to find a single community of closely related nodes for a given query node, but in general, nodes are possible to belong to several communities, and detecting all the potential communities for a given query node is much more challenging. In this work, we propose a novel approach called the centroid-based multiple local community detection (C-MLC) to find all the communities for a query node. Differing from the existing local community detection methods that directly find a community from the query node, we assume that every community contains a “centroid” node, which locates in the core of the community and can be used to identify the community. Then, a query node corresponds to several centroid nodes if the query node belongs to multiple communities. The key ideas of C-MLC are that C-MLC automatically determines the number of communities containing the query node by finding the related centroid nodes and uses each query node together with the centroid node to uncover the corresponding community based on a set of high-quality seeds. Through extensive evaluations on real-world networks and synthetic networks, C-MLC outperforms the state-of-the-art methods significantly, demonstrating that finding the centroid nodes is a better approach to uncover the multiple local communities. Boyu Li 0003, Dany Kamuhanda, Kun He 0001 |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2024 | LUSTC: A Novel Approach for Predicting Link States on Dynamic Attributed NetworksabstractPredicting link states on dynamic attributed networks is one of the most fundamental problems for network analysis. To accurately predict the formations and disappearances of links, we need to utilize three types of information, that is, structural information, node content information, and temporal information, simultaneously. To this end, we propose a novel approach called LUSTC for predicting links and unlinks with structural information, temporal information, and node content information on dynamic attributed networks. For each snapshot, LUSTC first randomly chooses some sets of active nodes whose degree changes between adjacent snapshots and collects higher-order structural information using a random walk method based on the active nodes. Then, LUSTC generates two global matrices containing the temporal structural information and temporal node content information, respectively, as well as a sequence of auxiliary matrices for reconstructing the structural information and node content information of snapshots. These generated matrices are optimized using a nonnegative matrix factorization (NMF) method based on the structural information and node content information of snapshots. Finally, LUSTC estimates the similarity matrix for future snapshots to predict the links that are more likely to be formed or broken. The experimental results on various dynamic attributed networks demonstrate the effectiveness of LUSTC on the link prediction and unlink prediction tasks. Christina Muro, Boyu Li 0003, Kun He 0001 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2024 | Spatio-temporal Contrastive Learning-enhanced GNNs for Session-based RecommendationabstractSession-based recommendation (SBR) systems aim to utilize the user’s short-term behavior sequence to predict the next item without the detailed user profile. Most recent works try to model the user preference by treating the sessions as between-item transition graphs and utilize various graph neural networks (GNNs) to encode the representations of pair-wise relations among items and their neighbors. Some of the existing GNN-based models mainly focus on aggregating information from the view of spatial graph structure, which ignores the temporal relations within neighbors of an item during message passing and the information loss results in a sub-optimal problem. Other works embrace this challenge by incorporating additional temporal information but lack sufficient interaction between the spatial and temporal patterns. To address this issue, inspired by the uniformity and alignment properties of contrastive learning techniques, we propose a novel framework called Session-based Recommendation with Spatio-temporal Contrastive Learning-enhanced GNNs (RESTC). The idea is to supplement the GNN-based main supervised recommendation task with the temporal representation via an auxiliary cross-view contrastive learning mechanism. Furthermore, a novel global collaborative filtering graph embedding is leveraged to enhance the spatial view in the main task. Extensive experiments demonstrate the significant performance of RESTC compared with the state-of-the-art baselines. We release our source code at https://github.com/SUSTechBruce/RESTC-Source-code . Zhongwei Wan, Xin Liu 0039, Benyou Wang, Jiezhong Qiu, Boyu Li 0003, Ting Guo 0005, Guangyong Chen, Yang Wang 0002 |
ACM Trans. Inf. Syst. | 5 |
| 2023 | ConGCN: Factorized Graph Convolutional Networks for Consensus Recommendation
Boyu Li 0003, Ting Guo 0005, Xingquan Zhu 0001, Yang Wang 0002, Fang Chen 0001 |
ECML/PKDD (4) | 1 |
| 2023 | SGCCL: Siamese Graph Contrastive Consensus Learning for Personalized RecommendationabstractContrastive-learning-based neural networks have recently been introduced to recommender systems, due to their unique advantage of injecting collaborative signals to model deep representations, and the self-supervision nature in the learning process. Existing contrastive learning methods for recommendations are mainly proposed through introducing augmentations to the user-item (U-I) bipartite graphs. Such a contrastive learning process, however, is susceptible to bias towards popular items and users, because higher-degree users/items are subject to more augmentations and their correlations are more captured. In this paper, we advocate a Siamese Graph Contrastive Consensus Learning (SGCCL) framework, to explore intrinsic correlations and alleviate the bias effects for personalized recommendation. Instead of augmenting original U-I networks, we introduce siamese graphs, which are homogeneous relations of user-user (U-U) similarity and item-item (I-I) correlations. A contrastive consensus optimization process is also adopted to learn effective features for user-item ratings, user-user similarity, and item-item correlation. Finally, we employ the self-supervised learning coupled with the siamese item-item/user-user graph relationships, which ensures unpopular users/items are well preserved in the embedding space. Different from existing studies, SGCCL performs well on both overall and debiasing recommendation tasks resulting in a balanced recommender. Experiments on four benchmark datasets demonstrate that SGCCL outperforms state-of-the-art methods with higher accuracy and greater long-tail item/user exposure. Boyu Li 0003, Ting Guo 0005, Xingquan Zhu 0001, Qian Li 0003, Yang Wang 0002, Fang Chen 0001 |
WSDM | 1 |
| 2023 | Self-Adaptive Predictive Passenger Flow Modeling for Large-Scale Railway SystemsabstractIntelligent transportation system (ITS) is a crucial symbol of smart cities, which aims to provide sustainable and efficient services to residents. Playing a vital role in ITS, railway systems have been integrating with multiple Internet of Things (IoT) devices to monitor real-time inbound passenger flows to ensure pedestrian safety. But it remains difficult to consolidate real-time information from different IoT sources and accurately estimate the future flow due to the coarse-grained data, potential impacts of dynamic interchanged passengers, and real-time predictive capability, which have greatly hindered the progress of ITS transformation in smart cities. To tackle these challenges, we propose a two-stage self-adaptive model for accurately and timely predicting passenger flow in metropolitan railway systems. In the first stage, a self-attention-based prediction model is introduced to predict the next-day passenger flow based on the historical boarding records captured by IoT devices. The proposed decomposing components transferring the discrete boarding records into continuous patterns enable the module to deliver a robust minute-level prediction. In the second stage, a real-time fine-tuning model is developed to adjust the predicted passenger flow based on real-time emergencies and short-term changes in passenger flows from IoT devices. The combination of an offline deep learning mechanism and a real-time reallocation algorithm ensures the real-time response without loss of accuracy. Our end-to-end framework has been deployed to the railway system in Greater Sydney Area, Australia, which can offer accurate predictions to trip planners for timetable design and provide timely decision support for controllers when emergencies happen. Boyu Li 0003, Ting Guo 0005, Yang Wang 0002, Amir Hossein Gandomi, Fang Chen 0001 |
IEEE Internet Things J. | 1 |
| 2023 | Link Prediction and Unlink Prediction on Dynamic NetworksabstractLink prediction on dynamic networks has been extensively studied and widely applied in various applications. However, existing methods only consider either the network structure or the temporal information, ignoring the potentialities of using both types of information together to comprehend the complex behaviors of dynamic networks. Moreover, temporal unlink prediction, which also plays an important role in the evolution of social networks, has not been paid much attention. Accurately predicting the links and unlinks on the future network greatly contributes to the network analysis that uncovers more latent relations between nodes. In this work, we assume that there are two kinds of relations between nodes, namely, long-term relations and short-term relations, and we propose an effective algorithm called LULS for temporal link prediction and unlink prediction based on such relations. Specifically, for each snapshot of a dynamic network, LULS first collects higher order structures as two topological matrices by applying short random walks. Then, LULS initializes and optimizes a global matrix and a sequence of temporary matrices for all the snapshots by using nonnegative matrix factorization (NMF) based on the topological matrices, where the global matrix denotes long-term relations and the temporary matrices represent short-term relations of snapshots. Finally, LULS calculates the similarity matrix of the future snapshot and predicts the links and unlinks for the future network. In addition, we further improve the prediction results by using graph regularization constraints to enhance the global matrix, resulting in that the global matrix contains a wealth of topological information and temporal information. The conducted experiments on real-world networks illustrate that LULS outperforms other baselines for both link prediction and unlink prediction tasks. Christina Muro, Boyu Li 0003, Kun He 0001 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2022 | A Two-Stage Self-adaptive Model for Passenger Flow Prediction on Schedule-Based Railway System
Boyu Li 0003, Ting Guo 0005, Yang Wang 0002, Amir Hossein Gandomi, Fang Chen 0001 |
PAKDD (3) | 1 |
| 2021 | Adaptive Graph Co-Attention Networks for Traffic Forecasting
Boyu Li 0003, Ting Guo 0005, Yang Wang 0002, Amir Hossein Gandomi, Fang Chen 0001 |
PAKDD (1) | 1 |
| 2021 | Local anatomy for personalised privacy protectionabstractAnonymisation technique has been extensively studied and widely applied for privacy-preserving data publishing. However, most existing methods ignore personal anonymity requirements. In these approaches, the microdata consist of three categories of attribute: explicit-identifier, quasi-identifier and sensitive attribute. In fact, the data sensitivity should be determined by individuals. An attribute is semi-sensitive if it contains both QI and sensitive values. In this paper, we propose a novel anonymisation approach, called local anatomy, to address personalised privacy protection. Local anatomy partitions the tuples who consider the value as sensitive into buckets inside each attribute. We conduct some experiments to illustrate that local anatomy can protect all the sensitive values and preserve great information utility. Additionally, we also present the concept of intelligent anonymisation system as our direction of future work. Boyu Li 0003, Yanheng Liu 0001, Minghai Wang, Geng Sun 0001 |
Int. J. Inf. Comput. Secur. | 1 |
| 2018 | Cross-Bucket Generalization for Information and Privacy PreservationabstractGeneralization is an effective technique for protecting confidential information of individuals, and has been studied by proposing numerous algorithms. However, the previous works do not separate the protection against identity disclosure and sensitive disclosure. Thus, when the requirement of attribute protection is higher than that of identity protection, generalization for l-diversity causes overprotection for identity and large mounts of information utility loss. This paper presents a novel approach, called cross-bucket generalization, as a solution to meet the problem. The rationale is to divide microdata into equivalence groups and buckets. First, it provides separate protection for identity and sensitive values, and the level of protection can be flexibly adjusted based on actual demands. Second, the sizes of equivalence groups and buckets are minimized as far as possible by only satisfying the protection requirements, which avoid the overprotection for identity and reduce information loss. The experiments we conducted illustrate the effectiveness of our solution. Boyu Li 0003, Yanheng Liu 0001, Xu Han 0005 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2015 | Reverse twin plant for efficient diagnosability testing and optimizing
Boyu Li 0003, Ting Guo 0005, Xingquan Zhu 0001, Zhanshan Li |
Eng. Appl. Artif. Intell. | 1 |