VLDB 2026 Research / reviewers in the wild / expert
Borui Cai
dblp:194/1410
· DBLP profile ↗
19ranked-venue papers
9as first author
17since 2021 · last 2025
0000-0002-1216-5166ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ContxE: Attention-based Context Aggregation for Temporal Knowledge Graph CompletionabstractKnowledge graph completion (KGC) methods aim to predict missing links by learning from existing facts in a knowledge graph. Different from KGC, Temporal Knowledge Graph Completion (TKGC) further incorporates the time validity of facts (tagged timestamps) during the learning and inference to improve the completion accuracy. Many TKGC methods achieve this by projecting the static entity representations (time-invariant) of KGC embedding methods to time-dependent representations, which vary across timestamps. However, when measuring a fact, these TKGC methods only consider its subject/object entity representations corresponding to the tagged timestamp, but ignore their historical contexts that normally carry essential supportive information. With this observation, we propose a novel context aggregation (ContxE) method to include historical contexts of subject/object entities for TKGC. To achieve that, we propose a linear-rotary time embedding to obtain time-dependent entity representations that can preserve temporal relationships, and a relation-based attention to aggregate historical context for the score measurement. Comprehensive experiments on three temporal knowledge graph datasets show that the proposed ContxE achieves improved knowledge graph completion results compared to strong counterpart methods. Borui Cai, Yong Xiang 0001, Longxiang Gao, Jiong Jin, Junfeng Wu 0010, Tom H. Luan |
IJCNN | 1 |
| 2025 | FedAdap: An Adaptive Federated Knowledge Graph Embedding Framework for Tackling KGs Heterogeneity via Partial Model SharingabstractKnowledge Graph Embedding (KGE) is a technique used to capture structural information from Knowledge Graphs (KGs), enabling various downstream applications such as recommender system. KGE models trained on integrated KGs from multiple organizations tend to outperform those trained on a single KG, owing to the greater richness and diversity of information. Therefore, Federated Knowledge Graph Embedding (FKGE) emerges as a promising approach for privacy-preserving training of KGE models on KGs across organizations (clients). Existing FKGE framework learns a uniform global KGE model that achieves global optima by minimizing aggregated loss across clients. However, heterogeneity among KGs often leads to divergent local optima. This presents a fundamental trade-off: ensuring global optima can compromise local performance, while focusing on local optima can decrease global model utility. To overcome this, we propose Federated Local Adaptive Knowledge Graph Embedding (FedAdap) by drawing inspiration from partial federated learning. FedAdap employs a multilayer convolutional neural network, wherein its lower layers are shared across clients to learn shared information, it maps a seed KGE model into an alignment vector space representation. Its upper layers remain private, transforming the alignment vector space representation to an adaptive KGE model tailored to the local KG. Through this, FedAdap allows clients to leverage shared information while maintaining local adaptability and mitigating the impact of KGs heterogeneity. Experiments on data sets FB15k-237 and NELL-995 show that FedAdap outperforms its counterparts in link prediction tasks. Borui Cai, Yong Xiang 0001, Keshav Sood |
IJCNN | 2 |
| 2025 | A federated compositional knowledge graph embedding for communication efficiencyabstractKnowledge Graph Embedding (KGE), which automatically capture structural information from Knowledge Graphs (KGs), are essential for enhancing various downstream tasks, such as recommender systems. To further improve the effectiveness of KGE models, Federated Knowledge Graph Embedding (FKGE) has been introduced, enabling the privacy-preserving integration of KGs across multiple organizations. However, existing FKGE frameworks require aggregation of a large global KGE model (embeddings). resulting in significant communication overhead, thereby reducing the efficiency and utility of FKGE in practical scenarios. To address this challenge, we propose Federated Compositional Knowledge Graph Embedding (FedComp), which enhances communication efficiency by leveraging the compositional characteristics of KG entities. In FedComp, we design a lightweight global model that represents shareable latent features of entities. These global latent features are composed into personalized KGE models with local embedding generators on the clients, improving both local adaptability and performance. By this, FedComp can significantly reduce the number of parameters that need to be transmitted Experimental results show that FedComp outperforms state-of-the-art FKGE frameworks on link prediction accuracy, with only around 1.0% communication overhead compared to counterpart frameworks. Borui Cai, Yong Xiang 0001, Yao Zhao 0006, Md Palash Uddin, Keshav Sood |
Knowl. Based Syst. | 2 |
| 2024 | ICPR 2024 Competition on Multi-line Mathematical Expressions Recognition
Tianrui Zong, Wei Luo 0001, Wenhao Yu 0013, Borui Cai, He Zhang 0034, Fenghong Liu 0002, Qinqin Yan, Liangcai Gao |
ICPR (34) | 5 |
| 2024 | Early Discovery of Key Innovative Publications by Analyzing Emerging Topic Trends
Junfeng Wu 0010, Xiangmin Zhou, Guangyan Huang, Borui Cai, Guang-Li Huang, Hui Zheng 0001, Chihung Chi, Jing He 0004 |
WISE (1) | 4 |
| 2024 | SE-shapelets: Semi-supervised Clustering of Time Series Using Representative ShapeletsabstractShapelets that discriminate time series using local features (subsequences) are promising for time series clustering. Existing time series clustering methods may fail to capture representative shapelets because they discover shapelets from a large pool of uninformative subsequences, and thus result in low clustering accuracy. This paper proposes a Semi-supervised Clustering of Time Series Using Representative Shapelets (SE-Shapelets) method, which utilizes a small number of labeled and propagated pseudo-labeled time series to help discover representative shapelets, thereby improving the clustering accuracy. In SE-Shapelets, we propose two techniques to discover representative shapelets for the effective clustering of time series. (1) A salient subsequence chain (SSC) that can extract salient subsequences (as candidate shapelets) of a labeled/pseudo-labeled time series, which helps remove massive uninformative subsequences from the pool. (2) A linear discriminant selection (LDS) algorithm to identify shapelets that can capture representative local features of time series in different classes, for convenient clustering. Experiments on UCR time series datasets demonstrate that SE-shapelets discovers representative shapelets and achieves higher clustering accuracy than counterpart semi-supervised time series clustering methods. Borui Cai, Guangyan Huang, Shuiqiao Yang, Yong Xiang 0001, Chihung Chi |
Expert Syst. Appl. | 1 |
| 2024 | From Wide to Deep: Dimension Lifting Network for Parameter-Efficient Knowledge Graph EmbeddingabstractKnowledge graph embedding (KGE) that maps entities and relations into vector representations is essential for downstream applications. Conventional KGE methods require high-dimensional representations to learn the complex structure of knowledge graph, but lead to oversized model parameters. Recent advances reduce parameters by low-dimensional entity representations, while developing techniques (e.g., knowledge distillation or reinvented representation forms) to compensate for reduced dimension. However, such operations introduce complicated computations and model designs that may not benefit large knowledge graphs. To seek a simple strategy to improve the parameter efficiency of conventional KGE models, we take inspiration from that deeper neural networks require exponentially fewer parameters to achieve expressiveness comparable to wider networks for compositional structures. We view all entity representations as a single-layer embedding network, and conventional KGE methods that adopt high-dimensional entity representations equal widening the embedding network to gain expressiveness. To achieve parameter efficiency, we instead propose a deeper embedding network for entity representations, i.e., a narrow entity embedding layer plus a multi-layer dimension lifting network (LiftNet). Experiments on three public datasets show that by integrating LiftNet, four conventional KGE methods with 16-dimensional representations achieve comparable link prediction accuracy as original models that adopt 512-dimensional representations, saving 68.4% to 96.9% parameters. Borui Cai, Yong Xiang 0001, Longxiang Gao, Di Wu 0050, He Zhang 0034, Jiong Jin, Tom H. Luan |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | ARFL: Adaptive and Robust Federated LearningabstractFederated Learning (FL) is a machine learning technique that enables multiple local clients holding individual datasets to collaboratively train a model, without exchanging the clients' datasets. Conventional FL approaches often assign a fixed workload (local epoch) and step size (learning rate) to the clients during the client-side local model training and utilize all collaborating trained models' parameters evenly during the server-side global model aggregation. Consequently, they frequently experience problems with data heterogeneity and high communication costs. In this paper, we propose a novel FL approach to mitigate the above problems. On the client side, we propose an adaptive model update approach that optimally allocates a needful number of local epochs and dynamically adjusts the learning rate to train the local model and regularizes the conventional objective function by adding a proximal term to it. On the server side, we propose a robust model aggregation strategy that potentially supplants the local outlier updates (models' weights) prior to the aggregation. We provide the theoretical convergence results and perform extensive experiments on different data setups over the MNIST, CIFAR-10, and Shakespeare datasets, which manifest that our FL scheme surpasses the baselines in terms of communication speedup, test-set performance, and global convergence. Md Palash Uddin, Yong Xiang 0001, Borui Cai, Xuequan Lu, John Yearwood, Longxiang Gao |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | Temporal Knowledge Graph Completion: A SurveyabstractKnowledge graph completion (KGC) predicts missing links and is crucial for real-life knowledge graphs, which widely suffer from incompleteness. KGC methods assume a knowledge graph is static, but that may lead to inaccurate prediction results because many facts in the knowledge graphs change over time. Emerging methods have recently shown improved prediction results by further incorporating the temporal validity of facts; namely, temporal knowledge graph completion (TKGC). With this temporal information, TKGC methods explicitly learn the dynamic evolution of the knowledge graph that KGC methods fail to capture. In this paper, for the first time, we comprehensively summarize the recent advances in TKGC research. First, we detail the background of TKGC, including the preliminary knowledge, benchmark datasets, and evaluation metrics. Then, we summarize existing TKGC methods based on how the temporal validity of facts is used to capture the temporal dynamics. Finally, we conclude the paper and present future research directions of TKGC. Borui Cai, Yong Xiang 0001, Longxiang Gao, He Zhang 0034, Jianxin Li 0001 |
IJCAI | 1 |
| 2023 | Hybrid variational autoencoder for time series forecastingabstractVariational autoencoders (VAE) are powerful generative models that learn the latent representations of input data as random variables. Recent studies show that VAE can flexibly learn the complex temporal dynamics of time series and achieve more promising forecasting results than deterministic models. However, a major limitation of existing works is that they fail to jointly learn the local patterns (e.g., seasonality and trend) and temporal dynamics of time series for forecasting. Accordingly, we propose a novel hybrid variational autoencoder (HyVAE) to integrate the learning of local patterns and temporal dynamics by variational inference for time series forecasting. Experimental results on four real-world datasets show that the proposed HyVAE achieves better forecasting results than various counterpart methods, as well as two HyVAE variants that only learn the local patterns or temporal dynamics of time series, respectively. Borui Cai, Shuiqiao Yang, Longxiang Gao, Yong Xiang 0001 |
Knowl. Based Syst. | 1 |
| 2023 | Variational co-embedding learning for attributed network clustering
Shuiqiao Yang, Sunny Verma, Borui Cai, Jiaojiao Jiang 0001, Kun Yu 0001, Fang Chen 0001, Shui Yu 0001 |
Knowl. Based Syst. | 3 |
| 2023 | CoSS: Leveraging Statement Semantics for Code SummarizationabstractAutomated code summarization tools allow generating descriptions for code snippets in natural language, which benefits software development and maintenance. Recent studies demonstrate that the quality of generated summaries can be improved by using additional code representations beyond token sequences. The majority of contemporary approaches mainly focus on extracting code syntactic and structural information from abstract syntax trees (ASTs). However, from the view of macro-structures, it is challenging to identify and capture semantically meaningful features due to fine-grained syntactic nodes involved in ASTs. To fill this gap, we investigate how to learn more code semantics and control flow features from the perspective of code statements. Accordingly, we propose a novel model entitled CoSS for code summarization. CoSS adopts a Transformer-based encoder and a graph attention network-based encoder to capture token-level and statement-level semantics from code token sequence and control flow graph, respectively. Then, after receiving two-level embeddings from encoders, a joint decoder with a multi-head attention mechanism predicts output sequences verbatim. Performance evaluations on Java, Python, and Solidity datasets validate that CoSS outperforms nine state-of-the-art (SOTA) neural code summarization models in effectiveness and is competitive in execution efficiency. Further, the ablation study reveals the contribution of each model component. Chaochen Shi, Borui Cai, Yao Zhao 0006, Longxiang Gao, Keshav Sood, Yong Xiang 0001 |
IEEE Trans. Software Eng. | 2 |
| 2022 | Overview of NLPCC2022 Shared Task 5 Track 2: Named Entity Recognition
Borui Cai, He Zhang 0034, Fenghong Liu 0002, Ming Liu 0028, Tianrui Zong |
NLPCC (2) | 1 |
| 2022 | Overview of NLPCC2022 Shared Task 5 Track 1: Multi-label Classification for Scientific Literature
Ming Liu 0028, He Zhang 0034, Yangjie Tian, Tianrui Zong, Borui Cai, Ruohua Xu |
NLPCC (2) | 5 |
| 2022 | Robust cross-network node classification via constrained graph mutual information
Shuiqiao Yang, Borui Cai, Taotao Cai, Jiaojiao Jiang 0001, Bing Li 0002, Jianxin Li 0001 |
Knowl. Based Syst. | 2 |
| 2022 | A Covert Electricity-Theft Cyberattack Against Machine Learning-Based Detection ModelsabstractA Covert Electricity-Theft Cyberattack Against Machine Learning-Based Detection Models Lei Cui 0006, Lei Guo 0019, Longxiang Gao, Borui Cai, Youyang Qu, Yipeng Zhou, Shui Yu 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2021 | An Accurate Negative Survey Using Answer Confidence LevelabstractNegative survey is an effective method to protect the privacy of survey participants and has many applications. Different from normal survey, it hides personal privacy by requiring an individual participant to provide a negative option as the answer to the survey question. Answers of participants can further be reconstructed into the distribution of positive options. However, the current negative survey that assumes all participants are 100 percent confident in their answers introduces some precision loss to the reconstruction process. In this paper, we provide an accurate negative survey by allowing participants to annotate their confidence levels to the answers. In particular, we develop a novel Negative Survey to Positive Survey with Answer Confidence Level (NStoPS-CL) algorithm to reconstruct the negative survey with answer confidence level and further increase the accuracy of reconstruction. Experiments demonstrate that NStoPS-CL reliably improves the reconstruction accuracy by testing answer confidence level under different conditions (i.e., dataset size, number of question options and dataset distributions), while balancing reconstruction accuracy and privacy well. Borui Cai, Guangyan Huang, Chihung Chi, Yanfeng Shu |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2020 | Clustering Hashtags Using Temporal Patterns
Borui Cai, Guangyan Huang, Shuiqiao Yang, Yong Xiang 0001, Chihung Chi |
WISE (1) | 1 |
| 2018 | Clustering of Multiple Density Peaks
Borui Cai, Guangyan Huang, Yong Xiang 0001, Jing He 0004, Guang-Li Huang, Xiangmin Zhou |
PAKDD (3) | 1 |