Silin Zhou

dblp:307/2139 · DBLP profile ↗
← Back
14ranked-venue papers
8as first author
14since 2021 · last 2026
0009-0006-2889-8862ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Region-Point Joint Representation for Effective Trajectory Similarity Learning
abstract
Recent learning-based methods have reduced the computational complexity of traditional trajectory similarity computation, but state-of-the-art (SOTA) methods still fail to leverage the comprehensive spectrum of trajectory information for similarity modeling. To tackle this problem, we propose RePo, a novel method that jointly encodes Region-wise and Point-wise features to capture both spatial context and fine-grained moving patterns. For region-wise representation, the GPS trajectories are first mapped to grid sequences, and spatial context are captured by structural features and semantic context enriched by visual features. For point-wise representation, three lightweight expert networks extract local, correlation, and continuous movement patterns from dense GPS sequences. Then, a router network adaptively fuses the learned point-wise features, which are subsequently combined with region-wise features using cross-attention to produce the final trajectory embedding. To train RePo, we adopt a contrastive loss with hard negative samples to provide similarity ranking supervision. Experiment results show that RePo achieves an average accuracy improvement of 22.2% over SOTA baselines across all evaluation metrics.
Silin Zhou, Lisi Chen 0001, Shuo Shang
AAAI2
2026 JanusMM: A Benchmark for Self-Deprecation Understanding in Real-World Multimodal Conversations
abstract
Xinyi Xu, Bingguang Hao, Yongyi Xiong, Zimo Chen, Xinchen Liu, Hongxin Guo, Xuelong Wang, Silin Zhou, Shihan Dou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Bingguang Hao, Yongyi Xiong, Zimo Chen, Xinchen Liu, Hongxin Guo, Xuelong Wang, Silin Zhou, Shihan Dou
ACL (1)8
2025 Building Efficient LLM Pipeline for Human Mobility Prediction
abstract
Human mobility prediction is a fundamental problem in spatio-temporal data mining with broad applications in urban computing and transportation systems. While large language models (LLMs) have demonstrated strong sequence modeling capabilities, directly adapting them to structured mobility data remains challenging due to long input sequences and efficiency limitations. In this study, we propose ELP-Mob, an efficient framework that reformulates mobility prediction as a language modeling problem. ELP-Mob employs an instruction-style prompt design that incorporates user mobility profiles, historical trajectories, and target future time slots, enabling LLMs to understand mobility patterns and make predictions. To further enhance efficiency, ELP-Mob includes a data selection strategy that reduces redundancy by sampling informative subsets of training users, and a dynamic splitting strategy with token-length control, which scales to long histories while reducing computational overhead. In the GISCUP 2025, ELP-Mob achieved 6th place on the official leaderboard. The source code is publicly available at https://github.com/chwang0721/ELP-Mob.
Chenhao Wang 0007, Silin Zhou, Lisi Chen 0001, Shuo Shang
SIGSPATIAL/GIS2
2025 Blurred Encoding for Trajectory Representation Learning
abstract
Trajectory representation learning (TRL) maps trajectories to vector embeddings and facilitates tasks such as trajectory classification and similarity search. State-of-the-art (SOTA) TRL methods transform raw GPS trajectories to grid or road trajectories to capture high-level travel semantics, i.e., regions and roads. However, they lose fine-grained spatial-temporal details as multiple GPS points are grouped into a single grid cell or road segment. To tackle this problem, we propose the BLU rred Encoding method, dubbed BLUE, which gradually reduces the precision of GPS coordinates to create hierarchical patches with multiple levels. The low-level patches are small and preserve fine-grained spatial-temporal details, while the high-level patches are large and capture overall travel patterns. To complement different patch levels with each other, our BLUE is an encoder-decoder model with a pyramid structure. At each patch level, a Transformer is used to learn the trajectory embedding at the current level, while pooling prepares inputs for the higher level in the encoder, and up-resolution provides guidance for the lower level in the decoder. BLUE is trained using the trajectory reconstruction task with the MSE loss. We compare BLUE with 8 SOTA TRL methods for 3 downstream tasks, the results show that BLUE consistently achieves higher accuracy than all baselines, outperforming the best-performing baselines by an average of 30.90%. Our code is available at https://github.com/slzhou-xy/BLUE.
Silin Zhou, Yao Chen 0008, Shuo Shang, Lisi Chen 0001, Bingsheng He, Ryosuke Shibasaki
KDD (2)1
2025 Grid and Road Expressions Are Complementary for Trajectory Representation Learning
abstract
Trajectory representation learning (TRL) maps trajectories to vectors that can be used for many downstream tasks. Existing TRL methods use either grid trajectories, capturing movement in free space, or road trajectories, capturing movement in a road network, as input. We observe that the two types of trajectories are complementary, providing either region and location information or providing road structure and movement regularity. Therefore, we propose a novel multimodal TRL method, dubbed GREEN, to jointly utilize Grid and Road trajectory Expressions for Effective representatioN learning. In particular, we transform raw GPS trajectories into both grid and road trajectories and tailor two encoders to capture their respective information. To align the two encoders such that they complement each other, we adopt a contrastive loss to encourage them to produce similar embeddings for the same raw trajectory and design a mask language model (MLM) loss to use grid trajectories to help reconstruct masked road trajectories. To learn the final trajectory representation, a dual-modal interactor is used to fuse the outputs of the two encoders via cross-attention. We compare GREEN with 7 state-of-the-art TRL methods for 3 downstream tasks, finding that GREEN consistently outperforms all baselines and improves the accuracy of the best-performing baseline by an average of 15.99%. Code and data are available at https://github.com/slzhou-xy/GREEN.
Silin Zhou, Shuo Shang, Lisi Chen 0001, Peng Han 0005, Christian S. Jensen
KDD (1)1
2025 Feature Enhanced Spatial-Temporal Trajectory Similarity Computation
abstract
Abstract Trajectory similarity computation is a fundamental function in many applications of urban data analysis, such as trajectory clustering, trajectory compression, and route planning. In this paper, we study trajectory similarity computation on the road network. However, existing methods have been designed primarily for road network trajectories with spatial information, while ignoring the important temporal information in the real world. To solve this problem, we propose a Feature Enhanced Spatial–Temporal trajectory similarity computation framework FEST, which is a graph neural network (GNN) and sequence model pipeline. We first use the GNN model to capture global information on the road network. In particular, we enhance the process with multi-graph to learn multiple signals from the road network on different aspects. In addition to the original road network topology signal, we also take into account the content signal to learn spatial–temporal features from trajectory traffic, as well as the adaptive similarity signal of the road network to learn hidden features. From these three signals, we construct a multi-graph and use GCN to learn road intersection embedding jointly. Next, we propose a feature-enhanced Transformer with spatial–temporal information to learn correlation within the trajectory, and we further use mean-pooling to get the final trajectory embedding. We compare FEST with six trajectory similarity computation methods on two real-world datasets. The results show that FEST consistently outperforms all baselines and can improve the accuracy of the best-performing baseline.
Silin Zhou, Chengrui Huang 0001, Yuntao Wen, Lisi Chen 0001
Data Sci. Eng.1
2024 RED: Effective Trajectory Representation Learning with Comprehensive Information
abstract
Trajectory representation learning (TRL) maps trajectories to vectors that can then be used for various downstream tasks, including trajectory similarity computation, trajectory classification, and travel-time estimation. However, existing TRL methods often produce vectors that, when used in downstream tasks, yield insufficiently accurate results. A key reason is that they fail to utilize the comprehensive information encompassed by trajectories. We propose a self-supervised TRL framework, called RED, which effectively exploits multiple types of trajectory information. Overall, RED adopts the Transformer as the backbone model and masks the constituting paths in trajectories to train a masked autoencoder (MAE). In particular, RED considers the moving patterns of trajectories by employing a R oad-aware masking strategy that retains key paths of trajectories during masking, thereby preserving crucial information of the trajectories. RED also adopts a spatial-temporal-user joint E mbedding scheme to encode comprehensive information when preparing the trajectories as model inputs. To conduct training, RED adopts D ual-objective task learning : the Transformer encoder predicts the next segment in a trajectory, and the Transformer decoder reconstructs the entire trajectory. RED also considers the spatial-temporal correlations of trajectories by modifying the attention mechanism of the Transformer. We compare RED with 9 state-of-the-art TRL methods for 4 downstream tasks on 3 real-world datasets, finding that RED can usually improve the accuracy of the best-performing baseline by over 5%.
Silin Zhou, Shuo Shang, Lisi Chen 0001, Christian S. Jensen, Panos Kalnis
Proc. VLDB Endow.1
2024 Cross-Task Multimodal Reinforcement for Long Tail Next POI Recommendation
abstract
Next Point-of-Interest (POI) recommendation seeks to recommend locations that users are most likely to visit next based on their historical trajectories, providing both users and service providers with substantial benefits. However, most next POI recommendation methods calculate the distances between POIs when mining spatial information and adjust their weights accordingly, ignoring the characteristics and multimedia content features of the regions in which POIs are located. In addition, the next POI recommendations suffer from the long tail effect, in which only a small portion of POIs appear frequently in users' recommendation lists due to their high popularity, while remainders maintain a low presence. To this end, we propose the cross-task multimodal reinforcement method which enriches the representations of regions by incorporating information from auxiliary domains. Moreover, we devise a cross-task reinforcement module to effectively integrate the local representations with pre-trained encoders from auxiliary domains. Actually, the enhanced region representations contain constructive district properties which are helpful to find proper POIs that suit users' tastes and thus alleviate the long tail effect. Experiments conducted on two real-world datasets indicate that our proposed method outperforms the state-of-the-art models in terms of both general performance and that of niche POIs.
Jiangfeng Du, Silin Zhou, Peng Han 0005, Shuo Shang
IEEE Trans. Multim.2
2023 GRLSTM: Trajectory Similarity Computation with Graph-Based Residual LSTM
abstract
The computation of trajectory similarity is a crucial task in many spatial data analysis applications. However, existing methods have been designed primarily for trajectories in Euclidean space, which overlooks the fact that real-world trajectories are often generated on road networks. This paper addresses this gap by proposing a novel framework, called GRLSTM (Graph-based Residual LSTM). To jointly capture the properties of trajectories and road networks, the proposed framework incorporates knowledge graph embedding (KGE), graph neural network (GNN), and the residual network into the multi-layer LSTM (Residual-LSTM). Specifically, the framework constructs a point knowledge graph to study the multi-relation of points, as points may belong to both the trajectory and the road network. KGE is introduced to learn point embeddings and relation embeddings to build the point fusion graph, while GNN is used to capture the topology structure information of the point fusion graph. Finally, Residual-LSTM is used to learn the trajectory embeddings.To further enhance the accuracy and robustness of the final trajectory embeddings, we introduce two new neighbor-based point loss functions, namely, graph-based point loss function and trajectory-based point loss function. The GRLSTM is evaluated using two real-world trajectory datasets, and the experimental results demonstrate that GRLSTM outperforms all the state-of-the-art methods significantly.
Silin Zhou, Jing Li 0034, Hao Wang 0005, Shuo Shang, Peng Han 0005
AAAI1
2023 Heterogeneous Region Embedding with Prompt Learning
abstract
The prevalence of region-based urban data has opened new possibilities for exploring correlations among regions to improve urban planning and smart-city solutions. Region embedding, which plays a critical role in this endeavor, faces significant challenges related to the varying nature of city data and the effectiveness of downstream applications. In this paper, we propose a novel framework, HREP (Heterogeneous Region Embedding with Prompt learning), which addresses both intra-region and inter-region correlations through two key modules: Heterogeneous Region Embedding (HRE) and prompt learning for different downstream tasks. The HRE module constructs a heterogeneous region graph based on three categories of data, capturing inter-region contexts such as human mobility and geographic neighbors, and intraregion contexts such as POI (Point-of-Interest) information. We use relation-aware graph embedding to learn region and relation embeddings of edge types, and introduce selfattention to capture global correlations among regions. Additionally, we develop an attention-based fusion module to integrate shared information among different types of correlations. To enhance the effectiveness of region embedding in downstream tasks, we incorporate prompt learning, specifically prefix-tuning, which guides the learning of downstream tasks and results in better prediction performance. Our experiment results on real-world datasets demonstrate that our proposed model outperforms state-of-the-art methods.
Silin Zhou, Lisi Chen 0001, Shuo Shang, Peng Han 0005
AAAI1
2023 Personalized Re-ranking for Recommendation with Mask Pretraining
abstract
Abstract Re-ranking is to refine the candidate ranking list of recommended items, such that the re-ranked list attracts users to purchase or click more items than the candidate one without re-ranking. Items in the candidate list are often ranked by their relevance to users’ interests. It is thus important to exploit the mutual influence between items in the re-ranking process. Existing re-ranking models focus on only the pairwise influence between two items, and have limited capability to exploit the local mutual influence in a group of items. Users often show successive interests on a group of relevant items, e.g., mobile phone, phone covers, wireless headset, namely scene. We propose a novel re-ranking model that jointly exploits the local mutual influence in scenes and the global mutual influence between different scenes. Scene representations are learned by GNN and multi-head attention, where GNN aims to learn local mutual influence while multi-head attention is to learn global mutual influence. To study the interaction between users and scenes, matrix factorization on users is utilized to obtain the user preference, which can be further applied to scenes to compute the scene scores. The final re-ranking list is generated by sorting the predicted scores of all scenes. To further mine user history information and item related user information, we also develop the extension pretraining module which relies on mask mechanism to support users and items high-quality embedding generation. We conduct a comprehensive evaluation on several real-world datasets. The experimental results demonstrate that our model substantially outperforms existing approaches.
Peng Han 0005, Silin Zhou, Zichen Xu 0001, Lisi Chen 0001, Shuo Shang
Data Sci. Eng.2
2023 Towards robust trajectory similarity computation: Representation-based spatio-temporal similarity quantification
Ke Li 0019, Silin Zhou, Lisi Chen 0001, Shuo Shang
World Wide Web (WWW)3
2023 Spatial-temporal fusion graph framework for trajectory similarity computation
Silin Zhou, Peng Han 0005, Di Yao 0001, Lisi Chen 0001, Xiangliang Zhang 0001
World Wide Web (WWW)1
2022 Multiple Behaviors Recommendation with Graph Learning
abstract
Recommendation systems have been extensively investigated by existing studies. However, previous work only targets a single type of user behavior (e.g., purchase) while ignoring the fact that users usually have multiple behaviors when browsing products (e.g., view, click, add-to-cart). These different behaviors can generate a large number of attributes about users. Meanwhile, in typical collaborative filtering (CF) systems, users and items are generally handled separately. Therefore, the associativity between users and items is not taken into account. To solve the above problems, we propose a novel framework named as Multiple Behaviors recommendation with Graph Learning (MBGL). To capture multiple characteristics of users and items, we construct a knowledge graph with multiple entity relations between users and items from a variety of data, such as purchase, view, and add-to-cart. We apply the knowledge graph embedding (KGE) method to learn the pre-training embedding of entity and relation. To better learn high-hop embedding from multi-behavior data, we construct a heterogeneous user-item graph and further design relation-aware GAT to learn graph embedding with relation type. To support high-order dependency in the graph, we introduce the residual network to solve feature smoothing and vanishing gradient problems. Moreover, to obtain the high-quality user and item embedding, we design an attention fusion layer to learn the fusion embedding and adopt the multi-task learning to predict users' preferences under different behaviors. Experiment results on real-world dataset show that our model outperforms other recommendation methods.
Silin Zhou, Lisi Chen 0001, Shuo Shang
MMSP1