Wujie Yan

dblp:309/6579 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0004-6084-6136ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Self-supervised Graph Neural Sequential Recommendation with Disentangling Long and Short-Term Interest
abstract
In real-world scenarios, a large amount of noise in user historical behaviors obstructs the reflection of their genuine interests. The long-tail distribution of user-item interactions also makes it difficult to capture interest evolution patterns from historical sequences. Moreover, as user behavior sequences continue to grow, solely relying on conventional sequence models is insufficient to extract user interest information and learn accurate sequence representations, thus limiting recommendation accuracy. To address these issues, we propose a self-supervised graph neural sequential recommendation model called LS4SRec, which disentangles users’ long- and short-term interests. Specifically, LS4SRec constructs two independent interest encoders to extract users’ long- and short-term interests. By utilizing the global user behavior sequence graph WITG to provide additional collaborative signals for each interaction sequence, we alleviate the issue of data sparsity. Subsequently, contrastive learning is applied to WITG to remove noise information and enhance the sequence representation. Further, interest allocation matrices and sequence models are utilized to model users’ interest evolution patterns. Finally, we introduce sequence graph data augmentation methods and long- and short-term interest pseudo-label construction methods to generate unsupervised signals that assist in model training. Extensive experiments conducted on real-world data validate the effectiveness of our proposed model. Our model implementation codes are available at the link https://github.com/jiubaoyibao/LS4SRec .
Liang He 0006, Wujie Yan, Tingzhou Yi, Heli Sun
Trans. Recomm. Syst.2
2025 Real-time Indexing for Large-scale Recommendation by Streaming Vector Quantization Retriever
abstract
Retrievers, which form one of the most important recommendation stages, are responsible for efficiently selecting possible positive samples to the later stages under strict latency limitations. Because of this, large-scale systems always rely on approximate calculations and indexes to roughly shrink candidate scale, with a simple ranking model. Most of the existing methods mainly focus on incorporating complicated ranking models. However, index structure is not improved, which also bottlenecks the whole effectiveness. In this paper, we propose a novel index structure: streaming Vector Quantization model, as a new generation of retrieval paradigm. Streaming VQ attaches items with indexes in real time, granting it immediacy. Moreover, through meticulous verification of possible variants, it achieves additional benefits like index balancing and reparability, enabling it to support complicated ranking models as existing approaches. Streaming VQ has been deployed and replaced all major retrievers in Douyin and Douyin Lite, resulting in remarkable user engagement gain.
Xingyan Bin, Jianfei Cui, Wujie Yan, Zhichen Zhao, Xintian Han, Chongyang Yan, Feng Zhang 0047, Zuotao Liu
KDD (2)3
2022 Graph Community Infomax
abstract
Graph representation learning aims at learning low-dimension representations for nodes in graphs, and has been proven very useful in several downstream tasks. In this article, we propose a new model, Graph Community Infomax (GCI), that can adversarial learn representations for nodes in attributed networks. Different from other adversarial network embedding models, which would assume that the data follow some prior distributions and generate fake examples, GCI utilizes the community information of networks, using nodes as positive(or real) examples and negative(or fake) examples at the same time. An autoencoder is applied to learn the embedding vectors for nodes and reconstruct the adjacency matrix, and a discriminator is used to maximize the mutual information between nodes and communities. Experiments on several real-world and synthetic networks have shown that GCI outperforms various network embedding methods on community detection tasks.
Heli Sun, Bing Lv, Wujie Yan, Liang He 0006, Shaojie Qiao
ACM Trans. Knowl. Discov. Data4