Chengrui Huang 0001

dblp:333/6384-1 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2026
0009-0004-8661-697XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2 · 2 since 2021
YearPublicationVenuePosition
2026 DPEPO: Diverse Parallel Exploration Policy Optimization for LLM-based Agents
abstract
JunShuo Zhang, Chengrui Huang, Feng Guo, Zihan Li, Ke Shi, Menghua Jiang, Jiguo Yu, Shuo Shang, Shen Gao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
JunShuo Zhang, Chengrui Huang 0001, Jiguo Yu, Shuo Shang, Shen Gao
ACL (1)2
2025 An Immersing Oriented Role-Playing Framework with Duplex Relationship Modeling
Yuntao Wen, Shen Gao, Chengrui Huang 0001, Yifan Wang 0023, Shuo Shang
DASFAA (6)3
2025 Generative Next POI Recommendation with Semantic ID
abstract
Point-of-interest (POI) recommendation systems aim to predict the next destinations of user based on their preferences and historical check-ins. Existing generative POI recommendation methods usually employ random numeric IDs for POIs, limiting the ability to model semantic relationships between similar locations. In this paper, we propose Generative Next POI Recommendation with Semantic ID (GNPR-SID), an LLM-based POI recommendation model with a novel semantic POI ID (SID) representation method that enhances the semantic understanding of POI modeling. There are two key components in our GNPR-SID: (1) a Semantic ID Construction module that generates semantically rich POI IDs based on semantic and collaborative features, and (2) a Generative POI Recommendation module that fine-tunes LLMs to predict the next POI using these semantic IDs. By incorporating user interaction patterns and POI semantic features into the semantic ID generation, our method improves the recommendation accuracy and generalization of the model. To construct semantically related SIDs, we propose a POI quantization method based on residual quantized variational autoencoder, which maps POIs into a discrete semantic space. We also propose a diversity loss to ensure that SIDs are uniformly distributed across the semantic space. Extensive experiments on three benchmark datasets demonstrate that GNPR-SID substantially outperforms state-of-the-art methods, achieving up to 16% improvement in recommendation accuracy.
Yuxi Huang 0005, Shen Gao, Yifan Wang 0023, Chengrui Huang 0001, Shuo Shang
KDD (2)5
2025 LLM-Based Agents for Tool Learning: A Survey
abstract
Abstract Human beings capable of making and using tools can accomplish tasks far beyond their innate abilities, and this paradigm of integration with tools may not be limited to humans themselves. Recently, the large language model (LLM) has demonstrated immense potential across various fields with its unique planning and reasoning abilities. However, there are still many challenges beyond its capabilities due to deficiencies in its training data and inherent illusions. Thus, integrating LLMs and tools into tool learning agents has become a new emerging research direction. To this end, we present a systematic investigation and comprehensive review of tool-learning agents in this paper. We start by introducing the definition of the tool learning task for Agents and then illustrating the typical architecture of the tool-learning models. Since these tools are all defined by users, LLM does not know what tools there are and what their functions are. Thus, LLMs should first find appropriate tools and split the tool retrieval methods into two categories: training-based and non-training-based. To accurately complete the user task, it is important to decompose the task into several sub-tasks and execute them in the correct order. Following that, we introduce the tool planning methods and organize these works by whether they rely on the model’s inherent reasoning capabilities for planning or utilize external reasoning tools. Due to the rapid development of this field, we also introduce an emerging frontier direction: using multimodal tools for LLM. In addition, we compile current open-source benchmarks and evaluation metrics, focusing on their scale, composition, calculation methods, and assessment dimensions. Next, we introduce several application scenarios for the LLM-based tool learning methods. Finally, we discuss the safety and ethical issues involved in tool learning.
Weikai Xu, Chengrui Huang 0001, Shen Gao, Shuo Shang
Data Sci. Eng.2
2025 Feature Enhanced Spatial-Temporal Trajectory Similarity Computation
abstract
Abstract Trajectory similarity computation is a fundamental function in many applications of urban data analysis, such as trajectory clustering, trajectory compression, and route planning. In this paper, we study trajectory similarity computation on the road network. However, existing methods have been designed primarily for road network trajectories with spatial information, while ignoring the important temporal information in the real world. To solve this problem, we propose a Feature Enhanced Spatial–Temporal trajectory similarity computation framework FEST, which is a graph neural network (GNN) and sequence model pipeline. We first use the GNN model to capture global information on the road network. In particular, we enhance the process with multi-graph to learn multiple signals from the road network on different aspects. In addition to the original road network topology signal, we also take into account the content signal to learn spatial–temporal features from trajectory traffic, as well as the adaptive similarity signal of the road network to learn hidden features. From these three signals, we construct a multi-graph and use GCN to learn road intersection embedding jointly. Next, we propose a feature-enhanced Transformer with spatial–temporal information to learn correlation within the trajectory, and we further use mean-pooling to get the final trajectory embedding. We compare FEST with six trajectory similarity computation methods on two real-world datasets. The results show that FEST consistently outperforms all baselines and can improve the accuracy of the best-performing baseline.
Silin Zhou, Chengrui Huang 0001, Yuntao Wen, Lisi Chen 0001
Data Sci. Eng.2
2025 Trajectory generation: a survey on methods and techniques
Chengrui Huang 0001, Chenhao Wang 0007, Lisi Chen 0001
GeoInformatica2
2025 Efficient Parallel Processing of Semantic Trajectory Similarity Joins
abstract
Matching similar pairs of trajectories, called trajectory similarity join, is a fundamental functionality for the Internet of Everything (IoE). We obverse that keyword-augmented trajectories are becoming increasingly popular. In this light, we investigate semantic trajectory similarity (STS) join that consists of two subproblems, threshold-based STS Join and top-k STS (k-STS) Join. Each semantic trajectory is a sequence of geo-textual objects with both location and text information. Specifically, given two sets of semantic trajectories and a threshold$\theta $or result number k, the STS Join returns all pairs of semantic trajectories from the two sets with spatio-textual similarity no less than$\theta $, and the k-STS Join returns k most similar pairs of semantic trajectories from the two sets. To enable efficient STS and k-STS Joins processing on large sets of semantic trajectories, we present a two-phase parallel search algorithm. We first group semantic trajectories based on their text information. The algorithm’s per-group searches are independent of each other and thus can be performed in parallel. We generate spatial and textual summaries for each trajectory batch and develop batch filtering techniques to prune unqualified trajectory pairs in a batch mode. Next, we propose a divide-and-conquer algorithm to derive bounds of spatial similarity and textual similarity between two semantic trajectories, which enable us filter out dissimilar trajectory pairs efficiently. Further, hierarchical batch filtering join algorithm is developed to process k-STS Join. Experimental study with large semantic trajectory data confirms that our algorithm of processing semantic trajectory join is capable of substantially outperforming well-designed baselines.
Shuo Shang, Chengrui Huang 0001, Lisi Chen 0001
IEEE Internet Things J.2
2025 An Efficient Parallel Mechanism for Processing Trajectory Split-and-Combine
abstract
With the increasing availability of time-dependent objects on Social internet of Things (SIoT), taking advantage of this data for SIoT-based route planning and recommendation is becoming increasingly imperative. To facilitate IoT-based route planning and recommendation, Trajectory Search by Locations (TSL) has been serving as a fundamental operation for data cleaning based on personalized requirements. However, existing TSL methods regard each trajectory as an indivisible object of time series. Because that the lengths and spatial distributions of trajectories may vary, traditional TRL methods often fail to return high-quality results, especially when the trajectory data is spare. Such limitation is considered to be a major bottleneck for improving the effectiveness of SIoT-based route planning and recommendation. To address the limitation, we propose a parallel mechanism for handling the problem of Trajectory Search by Locations through Parallel split-and-combine (TSL-Psc). The TSL-Psc problem is described as follows. Given a set of trajectories, a query sequence Q consisting of a sequence of timestamped locations, and a spatio-temporal similarity threshold θ, we retrieve a combined trajectory composed of sub-trajectories that (i) has similarity to Q no less than θ and (ii) contains the minimum number of sub-trajectory combinations. The resulting functionality of TSL-Psc targets a broad range of time-sensitive applications regarding SIoT, including traffic analysis, real-time group route planning, and ridesharing. To enable efficient TSL-Psc computation of trajectory data, we develop a three-phase Joint-Split Parallel Search (JSPS) algorithm on the basis of FSPS and GSPS. JSPS consists of pre-checking, split, and combine phases. We develop a spatial-first expansion algorithm and a temporal-first expansion algorithm that are capable of pruning search space in both spatial and temporal domains. Comprehensive experiments conducted on real-world datasets provide valuable insights into the algorithm’s performance, demonstrating that the proposed TSL-Psc problem formulation produces high-quality outcomes while the search algorithms deliver high-standard efficiency and scalability.
Shuo Shang, Chengrui Huang 0001, Xiaocheng Hu, Lisi Chen 0001
IEEE Internet Things J.2