Jianpeng Zhou

dblp:25/5237 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 LLM-augmented entity alignment: an unsupervised and training-free framework
abstract
Entity alignment (EA) is a fundamental task in knowledge graph (KG) integration, aiming to identify equivalent entities across different KGs for a unified and comprehensive representation. Recent advances have explored pre-trained language models (PLMs) to enhance the semantic understanding of entities, achieving notable improvements. However, existing methods face two major limitations. First, they rely heavily on human-annotated labels for training, leading to high computational costs and poor scalability. Second, some approaches use large language models (LLMs) to predict alignments in a multi-choice question format, but LLM outputs may deviate from expected formats, and predefined options may exclude correct matches, leading to suboptimal performance. To address these issues, we propose LEA, an LLM-augmented entity alignment framework that eliminates the need for labeled data and enhances robustness by mitigating information heterogeneity at both embedding and semantic levels. LEA first introduces an entity textualization module that transforms structural and textual information into a unified format, ensuring consistency and improving entity representations. It then leverages LLMs to enrich entity descriptions, enhancing semantic distinctiveness. Finally, these enriched descriptions are encoded into a shared embedding space, enabling efficient alignment through text retrieval techniques. To balance performance and computational cost, we further propose a selective augmentation strategy that prioritizes the most ambiguous entities for refinement. Experimental results on both homogeneous and heterogeneous KGs demonstrate that LEA outperforms existing models trained on 30 % labeled data, achieving a 30 % absolute improvement in Hit@1 score. As LLMs and text embedding models advance, LEA is expected to further enhance EA performance, providing a scalable and robust paradigm for practical applications. The code and dataset can be found at https://github.com/Longmeix/LEA.
Meixiu Long, Jiahai Wang, Junxiao Ma, Jianpeng Zhou, Siyuan Chen 0005
Neural Networks4
2025 Geometry-Guided Behavior Pattern Adaptation for Trajectory Prediction in Unseen Scenes
abstract
Pedestrian trajectory prediction aims to forecast future trajectories based on observed behaviors and surrounding conditions, and it is critical for applications like autonomous driving. Predicting trajectories in unseen scenes is challenging due to varying environments, elusive internal movement patterns, and complex social interactions. Existing methods face two limitations. Firstly, they struggle to effectively extract internal movement patterns from historical trajectories without labeled samples, which are often inaccessible in practice. Secondly, they fail to learn social interaction patterns across scenes, particularly when using angle-related features that are noise-sensitive and not strictly invariant to Euclidean transformations. To address these challenges, this paper introduces a Geometry-guided Behavior Pattern Adaptation (GBPA) method based on two geometric observations. Firstly, properly normalized historical trajectories are distributionally similar to full trajectories, allowing generation of pseudo-full trajectories for auxiliary training. Secondly, the discretized angular partitions, created by splitting the perceptive field into equal-sized fans, are invariant to Euclidean transformations and robust to noise. GBPA employs a test-time training strategy on scaled historical trajectories (T3SH) to adapt internal movement patterns without future trajectories and an angular partitioned attention (APA) mechanism to capture transferable social interaction patterns by differentiating neighbors’ effects. Experimental results on two datasets demonstrate that GBPA significantly improves prediction performance.
Yaqun Cui, Meixiu Long, Jinbiao Chen, Jianpeng Zhou, Jiahai Wang
IJCNN4
2025 Adaptive-solver framework for dynamic strategy selection in large language model reasoning
Jianpeng Zhou, Wanjun Zhong, Jiahai Wang
Inf. Process. Manag.1
2025 Question Embedding on Weighted Heterogeneous Information Network for Knowledge Tracing
abstract
Knowledge Tracing (KT) aims to predict students’ future performance on answering questions based on their historical exercise sequences. To alleviate the problem of data sparsity in KT, recent works have introduced auxiliary information to mine question similarity, resulting in the enhancement of question embeddings. Nonetheless, there remains a gap in developing an approach that effectively incorporates various forms of auxiliary information, including relational information (e.g., question–student , question–skill relation), relationship attributes (e.g., correctness indicating a student's performance on a question), and node attributes (e.g., student ability ). To tackle this challenge, the Similarity-enhanced Question Embedding (SimQE) method for KT is proposed, with its central feature being the utilization of weighted and attributed meta-paths for extracting question similarity. To capture multi-dimensional question similarity semantics by integrating multiple relations, various meta-paths are constructed for learning question embeddings separately. These embeddings, each encoding different similarity semantics, are then fused to serve the task of KT. To capture finer-grained similarity by leveraging the relationship attributes and node attributes on the meta-paths, the biased random walk algorithm is designed. In addition, the auxiliary node generation method is proposed to capture high-order question similarity. Finally, extensive experiments conducted on six datasets demonstrate that SimQE performs the best among 10 representative question embedding methods. Furthermore, SimQE proves to be more effective in alleviating the problem of data sparsity.
Shangheng Du, Jianpeng Zhou, Xiaoxuan Shen, Ruxia Liang
ACM Trans. Knowl. Discov. Data3
2021 Collaborative Embedding for Knowledge Tracing
Jianpeng Zhou, Kai Zhang 0038, Qing Li 0045, Zijian Lu
KSEM2
2020 Self-Adaptive Framework for Efficient Stream Data Classification on Storm
abstract
In this era of big data, stream data classification which is one of typical data stream applications has become more and more significant and challengeable. In these applications, it is obvious that data classification is much more frequent than model training. The ratio of stream data to be classified is rapid and time-varying, so it is an important problem to classify the stream data efficiently with high throughput. In this paper, we first analyze and categorize the current data stream machine learning algorithms according to their data structures. Then, we propose stream data classification topology (SDC-Topology) on Storm. For the classification algorithms based on the matrix, we propose self-adaptive stream data classification framework (SASDC-Framework) for efficient stream data classification on Storm. In SASDC-Framework, all the data sets arriving at the same unit time are partitioned into subsets with the nearly best partition size and processed in parallel. To select the nearly best partition size for the stream data sets efficiently, we adopt bisection method strategy and inverse distance weighted strategy. Extreme learning machine, which is a fast and accurate machine learning method based on matrix calculating, is used to test the efficiency of our proposals. According to evaluation results, the throughputs based on SASDC-Framework are 8-35 times higher than those based on SDC-Topology and the best throughput is more than 40000 prediction requests per second in our environment.
Shizhuo Deng, Shan Huang 0007, Chuncheng Yue, Jianpeng Zhou, Guoren Wang
IEEE Trans. Syst. Man Cybern. Syst.5
2018 Global Shuffle Grouping (GSG): A Load Balancing Strategy for Continuous Range Queries on Storm
abstract
Apache Storm is a distributed stream processing framework to support real-time processing of big data. Even if many stream grouping strategies have been implemented in Storm to partition stream data in order to maximize usability of resources, but they cannot efficiently support continuous range query. It is the basis of location based services, in which both queries and objects are moving. The reason is that the spatial semantics of the query (range and data distribution) cannot be expressed by those strategies, and this is easy to result in load imbalance. For this problem, we propose a load-balancing strategy called global shuffle grouping (GSG) to support efficient continuous range queries on Storm. There the cost of the query is estimated based on the range and density of moving objects. The continuous range queries are grouped according to their costs by the way of round-robin. For the queries belonging to the same group, they are distributed according to a counter array by another round-robin. Double round-robins ensure that the load distributions to multiple downstream bolts are balanced. We implemented continuous range query topology with GSG into Storm. Compared with the most practicable built-in grouping strategy shuffle grouping, our proposed grouping is able to reduce load imbalance degree and load standard deviation by 2-3 times and reduce load fluctuation by 1-2 times. The throughput can be improved up to nearly 20%.
Jianpeng Zhou, Hanhui Zhong
SERA3
2008 Developing a GIS-Based Information Management System for on-Site wastewater Treatment Facilities
abstract
On-site wastewater treatment facilities (WWTFs) collect, treat, and dispose wastewater from dwellings that are not connected to municipal wastewater collection and treatment systems. They serve about 25% of the total population in the United States from an estimated 26 million homes, businesses, and recreational facilities nationwide. There is currently no adequate coordinated information management system for on-site WWTFs. Given the increasing concern about environmental contamination and its effect on public health, it is necessary to provide a more adequate management tool for on-site WWTFs information. This paper presents the development of an integrated, GIS-based, on-site wastewater information management system, which includes three components: (1) a mobile GIS for field data collection; (2) a World Wide Web (WWW) interface for electronic submission of individual WWTF information to a centralized GIS database in a state department of public health or state environmental protection agency; and (3) a GIS for the display and management of on-site WWTFs information, along with other spatial information such as land use, soil types, streams, and topography. It is anticipated that this GIS-based on-site wastewater information management system will provide environmental protection agencies and public health organizations with a spatial framework for managing on-site WWTFs and assessing the risks related to surface discharges.
Shunfu Hu, Jianpeng Zhou
Int. J. Softw. Eng. Knowl. Eng.2