En Xu

dblp:222/6151 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
8since 2021 · last 2026
0000-0002-8654-2788ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 STEM: Structure-Tracing Evidence Mining for Knowledge Graphs-Driven Retrieval-Augmented Generation
abstract
Knowledge Graph-based Question Answering (KGQA) plays a pivotal role in complex reasoning tasks but remains constrained by two persistent challenges: the structural heterogeneity of Knowledge Graphs (KGs) often leads to semantic mismatch during retrieval, while existing reasoning path retrieval methods lack a global structural perspective.To address these issues, we propose Structure-Tracing Evidence Mining (STEM), a novel framework that reframes multi-hop reasoning as a schema-guided graph search task.First, we design a Semanticto-Structural Projection pipeline that leverages KG structural priors to decompose queries into atomic relational assertions and construct an adaptive query schema graph.Subsequently, we execute globally-aware node anchoring and subgraph retrieval to obtain the final evidence reasoning graph from KG.To more effectively integrate global structural information during the graph construction process, we design a Triple-Dependent GNN (Triple-GNN) to generate a Global Guidance Subgraph (Guidance Graph) that guides the construction.STEM significantly improves both the accuracy and evidence completeness of multi-hop reasoning graph retrieval, and achieves State-of-the-Art performance on multiple multi-hop benchmarks.Our source code is available at https: //github.com/PennyYu123/STEM_RAG.
En Xu, Haibiao Chen, Yinfei Xu
ACL (1)2
2026 A Diffusive Data Augmentation Framework for Reconstruction of Complex Network Evolutionary History
abstract
The evolutionary dynamics of complex systems encode critical information about their functional organization. In particular, the generation times of edges reveal key aspects of historical development in networked systems such as protein-protein interaction networks, ecosystems, and social networks. Accurately recovering these temporal processes is of significant scientific value-for example, in elucidating the mechanisms underlying protein interaction evolution. However, existing methods typically assume access to partially time-stamped networks and often struggle to generalize across domains. They perform poorly in recovering edge generation times in static networks without temporal annotations. To address this challenge, we propose a comparative paradigm that enables cross-network learning by jointly training on multiple temporal networks. This framework captures structural-temporal correlations that generalize across networks and improves accuracy by 16.98% on average compared to separate training strategies. Furthermore, to mitigate the scarcity of real temporal data, we introduce a novel diffusion-based generative model for producing Augmented Temporal Networks (ATNs) . By integrating both real and generated samples during training, our joint strategy yields an additional 5.46% improvement in predictive accuracy, demonstrating the effectiveness of data augmentation in enhancing generalization.
En Xu, Can Rong, Jingtao Ding, Yong Li 0008
IEEE Trans. Knowl. Data Eng.1
2025 Upper bound on the predictability of rating prediction in recommender systems
En Xu, Zhiwen Yu 0001, Hui Wang 0011, Helei Cui, Yunji Liang, Bin Guo 0001
Inf. Process. Manag.1
2024 Limits of predictability in top-N recommendation
En Xu, Zhiwen Yu 0001, Ying Zhang 0047, Bin Guo 0001, Lina Yao 0001
Inf. Process. Manag.1
2023 Quantifying predictability of sequential recommendation via logical constraints
En Xu, Zhiwen Yu 0001, Helei Cui, Lina Yao 0001, Bin Guo 0001
Frontiers Comput. Sci.1
2023 Modeling Within-Basket Auxiliary Item Recommendation with Matchability and Ubiquity
abstract
Within-basket recommendation is to recommend suitable items for the current basket with some already known items. The within-basket auxiliary item recommendation ( WBAIR ) is to recommend auxiliary items based on the primary items in the basket. Such a task exists in many real-life scenarios. Unlike the associations between items that can be transmitted in both directions, primary and auxiliary relationships are unidirectional. Then, the suitable matching patterns between primary and auxiliary items cannot be explored by traditional directionless methods. Therefore, we design the Matc4Rec algorithm to integrate the primary and auxiliary factors, and finally recommend items that not only match the interests of users but also satisfy the primary and auxiliary relationships between items. Specifically, we capture the pattern from three aspects: matchability within-basket , matchability between baskets , and ubiquity . By exploiting this pattern, the designed algorithm not only achieves good results on real-world datasets but also improves the interpretability of recommendations. As a result, we can know which commodities are suitable as auxiliary items. The experiment results demonstrate that our algorithm can also alleviate the cold start problem.
En Xu, Zhiwen Yu 0001, Zhuo Sun 0002, Bin Guo 0001, Lina Yao 0001
ACM Trans. Intell. Syst. Technol.1
2022 Transfer how much: a fine-grained measure of the knowledge transferability of user behavior sequences in social network
Bin Guo 0001, Yan Liu 0045, Yasan Ding, En Xu, Lina Yao 0001, Zhiwen Yu 0001
Data Min. Knowl. Discov.5
2021 Core Interest Network for Click-Through Rate Prediction
abstract
In modern online advertising systems, the click-through rate (CTR) is an important index to measure the popularity of an item. It refers to the ratio of users who click on a specific advertisement to the number of total users who view it. Predicting the CTR of an item in advance can improve the accuracy of the advertisement recommendation. And it is commonly calculated based on users’ interests. Thus, extracting users’ interests is of great importance in CTR prediction tasks. In the literature, a lot of studies treat the interaction between users and items as sequential data and apply the recurrent neural network (RNN) model to extract users’ interests. However, these solutions cannot handle the case when the sequence length is relatively long, e.g., over 100. This is because of the vanishing gradient problem of RNN, i.e., the model cannot learn a users’ previous behaviors that are too far away from the current moment. To address this problem, we propose a new Core Interest Network (CIN) model to mitigate the problem of a long sequence in the CTR prediction task with sequential data. In brief, we first extract the core interests of users and then use the refined data as the input of subsequent learning tasks. Extensive evaluations on real dataset show that our CIN model can outperform the state-of-the-art solutions in terms of prediction accuracy.
En Xu, Zhiwen Yu 0001, Bin Guo 0001, Helei Cui
ACM Trans. Knowl. Discov. Data1
2019 Inferring User Profile Attributes From Multidimensional Mobile Phone Sensory Data
abstract
User profile can be used to characterize a person and help us better understand him/her, which further can be utilized to provide enhanced personalized services. When using mobile phone, some of one's information are unavoidably and unobtrusively passed or stored, which makes it possible to draw the user profile. In this paper, we propose to infer user profile, including age, gender, and personality traits based on mobile phone sensory data. Specifically, we capture data when unlocking screen, playing games as well as some basic mobile phone information, app usage, and screen status by using common available sensors in commodity mobile phones. By analyzing the differences in users' phone usage, we extracted features for user profile inference. Random Forest regression and random forest classification models are separately used to estimate age and gender of the user while support vector regression algorithm is applied to identify personality traits. In addition, we evaluate the model through real-life experiments conducted with a total of 84 phone users. Experimental results show that our approach effective, achieving an RSME of 4.3696 in age estimation and precision of 91.70% in gender detection. As for personality traits identification, the root mean square errors of openness, conscientiousness, extraversion, agreeableness, and neuroticism are 0.29, 0.3506, 0.465, 0.3022, and 0.452, respectively.
Zhiwen Yu 0001, En Xu, He Du, Bin Guo 0001, Lina Yao 0001
IEEE Internet Things J.2