Jingyun Zhang 0001

dblp:127/3055-1 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0003-7215-0040ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SECTOR: structural entropy-based learning of spatiotemporal organisation in spatial transcriptomics
abstract
MOTIVATION: Spatial transcriptomics (ST) profiles gene expression in tissue context, enabling spatial domain detection. However, relatively few methods jointly recover discrete spatial domains and continuous within-section pseudotemporal trends in a single framework. Current spatiotemporal approaches often emphasise trajectory continuity to recover smooth progression-associated gradients, but this may blur neighbouring domain boundaries and reduce clustering accuracy. Conversely, specialised spatial clustering algorithms typically rely on external single-cell trajectory tools rather than providing an integrated, spatially aware pseudotime model. RESULTS: We introduce SECTOR (Structural Entropy-based Clustering and pseudoTime ORdering), a lightweight deep graph learning framework that unifies spatial domain detection and pseudotime inference. SECTOR optimises a differentiable structural entropy (SE) objective on a fused spatial-expression graph, with spatial total variation regularisation to promote tissue continuity. Across seven benchmark datasets spanning standard and modern high-resolution ST platforms, SECTOR consistently outperformed existing spatiotemporal methods in clustering accuracy and matched or exceeded leading spatial clustering algorithms, while maintaining modest computational demands. In human breast cancer and mouse olfactory bulb case studies, SECTOR recovered spatially organised pseudotime patterns supported by semivariance, transition-gene, enrichment and marker-gene analyses. Together, these results show that SE-based learning provides an effective and scalable strategy for modelling within-section spatiotemporal organisation in ST. AVAILABILITY: SECTOR is available on GitHub at https://github.com/lhbcb/SECTOR and archived on Figshare at https://doi.org/10.6084/m9.figshare.32029830.
Jingyun Zhang 0001, Weikang Gong, Guangjie Zeng, Hao Peng 0001
Bioinform.2
2026 Enhanced Pre-training for Recommendation via Hypergraph Structural Entropy
abstract
Research on recommender systems plays a crucial role in alleviating information overload amid the current proliferation of data while diminishing user decision-making and transaction costs within intricate environments. The prevailing recommendation models currently rely on graph-based methods, such as GCN, GAT, HGNN, and so on, which are constrained by the sparsity of training data and the underutilization of graph structures. In this work, we present EPRHSE, an E nhanced P re-training framework for R ecommendation based on H ypergraph S tructural E ntropy, which encodes the topology of the recommender system. We begin by designing two forms of pre-training tasks to capture the heterogeneous relationships among users or items. These pre-training tasks build multiple auxiliary task hypergraphs, compensate for the sparse interactions between users and items, and unveil latent information. Secondly, we introduce a new method for optimizing the hypergraph structure entropy. The method involves converting the hyperedge information in the hypergraph to form a high-dimensional encoding tree. Hypergraph structure entropy helps decode the essential structure of the recommendation bipartite graph and enables hierarchical clustering of users or items. Thirdly, we propose a hypergraph pooling training methodology incorporating pooling and unpooling layers into the hypergraph convolutional network to amalgamate high-order information. By transferring advanced community insights to primary users or items, the process of social diffusion is enhanced, consequently refining node embedding quality. Compared with 13 representative recommendation approaches on five real datasets, comprehensive experiments demonstrate the effectiveness and advantages of EPRHSE.
Jingyun Zhang 0001, Hao Peng 0001, Mingdai Yang, Philip S. Yu
ACM Trans. Inf. Syst.1
2025 Unsupervised Graph Clustering with Deep Structural Entropy
abstract
Research on Graph Structure Learning (GSL) provides key insights for graph-based clustering, yet current methods like Graph Neural Networks (GNNs), Graph Attention Networks (GATs), and contrastive learning often rely heavily on the original graph structure. Their performance deteriorates when the original graph's adjacency matrix is too sparse or contains noisy edges unrelated to clustering. Moreover, these methods depend on learning node embeddings and using traditional techniques like k-means to form clusters, which may not fully capture the underlying graph structure between nodes. To address these limitations, this paper introduces DeSE, a novel unsupervised graph clustering framework incorporating Deep Structural Entropy. It enhances the original graph with quantified structural information and deep neural networks to form clusters. Specifically, we first propose a method for calculating structural entropy with soft assignment, which quantifies structure in a differentiable form. Next, we design a Structural Learning layer (SLL) to generate an attributed graph from the original feature data, serving as a target to enhance and optimize the original structural graph, thereby mitigating the issue of sparse connections between graph nodes. Finally, our clustering assignment method (ASS), based on GNNs, learns node embeddings and a soft assignment matrix to cluster on the enhanced graph. The ASS layer can be stacked to meet downstream task requirements, minimizing structural entropy for stable clustering and maximizing node consistency with edge-based cross-entropy loss. Extensive comparative experiments are conducted on four benchmark datasets against eight representative unsupervised graph clustering baselines, demonstrating the superiority of the DeSE in both effectiveness and interpretability.
Jingyun Zhang 0001, Hao Peng 0001, Li Sun 0008, Guanlin Wu, Zhengtao Yu 0001
KDD (2)1
2025 Relational Prompt-Based Pre-Trained Language Models for Social Event Detection
abstract
Social Event Detection (SED) aims to identify significant events from social streams, and has a wide application ranging from public opinion analysis to risk management. In recent years, Graph Neural Network (GNN) based solutions have achieved state-of-the-art performance. However, GNN-based methods often struggle with missing and noisy edges between messages, affecting the quality of learned message embedding. Moreover, these methods statically initialize node embedding before training, which, in turn, limits the ability to learn from message texts and relations simultaneously. In this article, we approach social event detection from a new perspective based on Pre-trained Language Models (PLMs), and present \(\mathrm{RPLM}_{SED}\) ( R elational prompt-based P re-trained L anguage M odels for S ocial E vent D etection). We first propose a new pairwise message modeling strategy to construct social messages into message pairs with multi-relational sequences. Secondly, a new multi-relational prompt-based pairwise message learning mechanism is proposed to learn more comprehensive message representation from message pairs with multi-relational prompts using PLMs. Thirdly, we design a new clustering constraint to optimize the encoding process by enhancing intra-cluster compactness and inter-cluster dispersion, making the message representation more distinguishable. We evaluate the \(\mathrm{RPLM}_{SED}\) on three real-world datasets, demonstrating that the \(\mathrm{RPLM}_{SED}\) model achieves state-of-the-art performance in offline, online, low-resource, and long-tail distribution scenarios for social event detection tasks.
Hao Peng 0001, Yantuan Xian, Linqin Wang, Li Sun 0008, Jingyun Zhang 0001, Philip S. Yu
ACM Trans. Inf. Syst.7
2024 Multivariate Time-Series Anomaly Detection based on Enhancing Graph Attention Networks with Topological Analysis
abstract
Unsupervised anomaly detection in time series is essential in industrial applications, as it significantly reduces the need for manual intervention. Multivariate time series pose a complex challenge due to their feature and temporal dimensions. Traditional methods use Graph Neural Networks (GNNs) or Transformers to analyze spatial while RNNs to model temporal dependencies. These methods focus narrowly on one dimension or engage in coarse-grained feature extraction, which can be inadequate for large datasets characterized by intricate relationships and dynamic changes. This paper introduces a novel temporal model built on an enhanced Graph Attention Network (GAT) for multivariate time series anomaly detection called TopoGDN. Our model analyzes both time and feature dimensions from a fine-grained perspective. First, we introduce a multi-scale temporal convolution module to extract detailed temporal features. Additionally, we present an augmented GAT to manage complex inter-feature dependencies, which incorporates graph topology into node features across multiple scales, a versatile, plug-and-play enhancement that significantly boosts the performance of GAT. Our experimental results confirm that our approach surpasses the baseline models on four datasets, demonstrating its potential for widespread application in fields requiring robust anomaly detection. The code is available at https://github.com/ljj-cyber/TopoGDN.
Zhe Liu 0004, Jingyun Zhang 0001, Zhifeng Hao 0005, Li Sun 0008, Hao Peng 0001
CIKM3
2024 Unsupervised Social Bot Detection via Structural Information Theory
abstract
Research on social bot detection plays a crucial role in maintaining the order and reliability of information dissemination while increasing trust in social interactions. The current mainstream social bot detection models rely on black-box neural network technology, for example, Graph Neural Network, Transformer, and so on, which lacks interpretability. In this work, we present UnDBot, a novel unsupervised, interpretable, yet effective, and practical framework for detecting social bots. This framework is built upon structural information theory. We begin by designing three social relationship metrics that capture various aspects of social bot behaviors: posting type distribution , posting influence , and follow-to-follower ratio . Three new relationships are utilized to construct a new, unified, and weighted social multi-relational graph, aiming to model the relevance of social user behaviors and discover long-distance correlations between users. Second, we introduce a novel method for optimizing heterogeneous structural entropy. This method involves the personalized aggregation of edge information from the social multi-relational graph to generate a two-dimensional encoding tree. The heterogeneous structural entropy facilitates decoding of the substantial structure of the social bots network and enables hierarchical clustering of social bots. Third, a new community labeling method is presented to distinguish social bot communities by computing the user’s stationary distribution, measuring user contributions to network structure, and counting the intensity of user aggregation within the community. Compared with 10 representative social bot detection approaches, comprehensive experiments demonstrate the advantages of effectiveness and interpretability of UnDBot on 4 real social network datasets.
Hao Peng 0001, Jingyun Zhang 0001, Zhifeng Hao 0005, Angsheng Li, Zhengtao Yu 0001, Philip S. Yu
ACM Trans. Inf. Syst.2