VLDB 2026 Research / reviewers in the wild / expert
Jinsong Chen 0002
dblp:14/7450-2
· DBLP profile ↗
13ranked-venue papers
9as first author
13since 2021 · last 2026
0000-0001-7588-6713ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Signgt: signed attention-based graph transformer for graph representation learning
Jinsong Chen 0002, Gaichao Li, John E. Hopcroft, Kun He 0001 |
Knowl. Inf. Syst. | 1 |
| 2026 | Integrating deep clustering and multi-view graph neural networks for recommender system
Jiaxuan Song, Duantengchuan Li, Rui Zhang 0017, Jinsong Chen 0002 |
Knowl. Based Syst. | 7 |
| 2026 | NTFormer: A Composite Node Tokenized Graph Transformer for Node ClassificationabstractTokenized graph Transformers have advanced node classification by transforming graphs into token sequences, but existing methods suffer from limited flexibility due to single-type token generation, which captures partial graph information and requires tailored modifications. To address this, we propose NTFormer, a novel graph Transformer with a dedicated token generator called Node2Par. Node2Par constructs diverse token sequences for each node using multiple token elements (i.e., neighborhood tokens and node tokens) from both topology view and attribute view, enabling comprehensive expression of graph features from multi-perspectives. Leveraging the outputs of Node2Pars, NTFormer adopts a standard Transformer backbone without additional graph-aware modules and a learnable information fusion strategy to adaptively learn expressive node representations from generated different token sequences, eliminating the need for tailored encoding strategies. Extensive experiments on benchmark datasets including homophily and heterophily graphs showcase that NTFormer outperforms representative graph Transformers and GNNs in node classification. Jinsong Chen 0002, Siyu Jiang, Kun He 0001 |
IEEE Trans. Big Data | 1 |
| 2026 | Tokenized Heterogeneous Graph Transformer with Enhanced Local and Global Representation LearningabstractGraph Transformers have demonstrated superiority in handling complex heterogeneous graphs in recent years. However, existing models still face several inherent challenges: (1) reliance on manually designed meta-paths to encode explicit local graph heterogeneity; (2) inability to capture fine-grained global information from distant yet relevant nodes. To address these limitations, we introduce THFormer, a novel node tokenized heterogeneous graph Transformer that learns expressive node representations by incorporating local and global perspectives. From the local perspective, we employ multiple subsequences for different heterogeneous types to explicitly encode local semantic relations, eliminating the need for manually designed meta-paths. From the global perspective, we design a local masking and global sampling mechanism to construct global structural (semantic) sequences, effectively capturing fine-grained global structural (semantic) information. Subsequently, THFormer separately feeds the resulting global and local sequences into standard Transformer layers as model inputs. Since these sequences represent two distinct views of the same target node, their corresponding outputs are naturally aligned to generate self-supervisory signals for model training, further enhancing the expressiveness and reliability of the target node representation. Extensive experiments are conducted to validate the efficacy of THFormer, and the quantitative performance gains are 0.24%, 0.31%, 0.51%, and 0.81% on DBLP, ACM, IMDB, and Freebase, respectively. The experimental results demonstrate the superiority of THFormer over representative heterogeneous graph neural networks and graph Transformer models. Gaichao Li, Jinsong Chen 0002, Yangzhe Peng, Kun He 0001 |
ACM Trans. Knowl. Discov. Data | 2 |
| 2025 | Rethinking Tokenized Graph Transformers for Node ClassificationabstractNode tokenized graph Transformers (GTs) have shown promising performance in node classification. The generation of token sequences is the key module in existing tokenized GTs which transforms the input graph into token sequences, facilitating the node representation learning via Transformer. In this paper, we observe that the generations of token sequences in existing GTs only focus on the first-order neighbors on the constructed similarity graphs, which leads to the limited usage of nodes to generate diverse token sequences, further restricting the potential of tokenized GTs for node classification. To this end, we propose a new method termed SwapGT. SwapGT first introduces a novel token swapping operation based on the characteristics of token sequences that fully leverages the semantic relevance of nodes to generate more informative token sequences. Then, SwapGT leverages a Transformer-based backbone to learn node representations from the generated token sequences. Moreover, SwapGT develops a center alignment loss to constrain the representation learning from multiple token sequences, further enhancing the model performance. Extensive empirical results on various datasets showcase the superiority of SwapGT for node classification.
Code is available at https://github.com/JHL-HUST/SwapGT. Jinsong Chen 0002, Gaichao Li, John E. Hopcroft, Kun He 0001 |
NeurIPS | 1 |
| 2025 | Hybrid long-range dependency-aware graph convolutional network for node classification
Jinsong Chen 0002, Meng Wang 0039, Kun He 0001 |
Knowl. Inf. Syst. | 1 |
| 2025 | NAGphormer+: A Tokenized Graph Transformer With Neighborhood Augmentation for Node Classification in Large GraphsabstractGraph Transformers, emerging as a new architecture for graph representation learning, suffer from the quadratic complexity and can only handle graphs with at most thousands of nodes. To this end, we propose a Neighborhood Aggregation Graph Transformer (NAGphormer) that treats each node as a sequence containing a series of tokens constructed by our proposed Hop2Token module. For each node, Hop2Token aggregates the neighborhood features from different hops into different representations, producing a sequence of token vectors as one input. In this way, NAGphormer could be trained in a mini-batch manner and thus could scale to large graphs with millions of nodes. To further enhance the model's generalization, we propose NAGphormer+, an extended model of NAGphormer with a novel data augmentation method called Neighborhood Augmentation (NrAug). Based on the output of Hop2Token, NrAug simultaneously augments the features of neighborhoods from global as well as local views. In this way, NAGphormer+ can fully utilize the neighborhood information of multiple nodes, thereby undergoing more comprehensive training and improving the model's generalization capability. Extensive experiments on benchmark datasets from small to large demonstrate the superiority of NAGphormer+ against existing graph Transformers and mainstream GNNs, as well as the original NAGphormer. Jinsong Chen 0002, Kaiyuan Gao, Gaichao Li, Kun He 0001 |
IEEE Trans. Big Data | 1 |
| 2025 | GTPool: Graph Transformer Pooling With Diverse SamplingabstractGraph pooling techniques have emerged as powerful tools for downsampling graphs, yielding impressive results on various graph-level tasks like graph classification and generation. Node dropping pooling stands out as a significant approach in this domain, utilizing learnable scoring functions to drop nodes with relatively lower significance. However, previous node dropping methods face two critical limitations: (1) they often struggle to capture long-range dependencies effectively for each pooled node, potentially limiting the expressiveness of node representations. (2) by exclusively retaining the highest-scoring nodes, they tend to preserve similar nodes, thereby discarding valuable information residing in low-scoring nodes; To address these issues, we propose a novel Graph Transformer Pooling method termed GTPool. GTPool constructs a node dropping pooling layer based on the Transformer architecture, aiming to efficiently capture long-range pairwise interactions while promoting diverse sampling. A key component of GTPool is its scoring module, grounded in the self-attention mechanism, which assesses the relevance between node representations and the graph representation. This nuanced approach offers a more natural means of measuring node importance. Additionally, GTPool adopts Roulette Wheel Sampling (RWS), a diversified sampling method that is capable of preserving nodes across various scoring intervals, rather than solely focusing on higher-scoring nodes. Consequently, GTPool can effectively capture long-range information and identify more representative nodes, thereby outperforming existing popular graph pooling methods across 14 benchmark datasets. Gaichao Li, Jinsong Chen 0002, Kun He 0001 |
IEEE Trans. Big Data | 2 |
| 2024 | Leveraging Contrastive Learning for Enhanced Node Representations in Tokenized Graph TransformersabstractWhile tokenized graph Transformers have demonstrated strong performance in node classification tasks, their reliance on a limited subset of nodes with high similarity scores for constructing token sequences overlooks valuable information from other nodes, hindering their ability to fully harness graph information for learning optimal node representations. To address this limitation, we propose a novel graph Transformer called GCFormer. Unlike previous approaches, GCFormer develops a hybrid token generator to create two types of token sequences, positive and negative, to capture diverse graph information. And a tailored Transformer-based backbone is adopted to learn meaningful node representations from these generated token sequences. Additionally, GCFormer introduces contrastive learning to extract valuable information from both positive and negative token sequences, enhancing the quality of learned node representations. Extensive experimental results across various datasets, including homophily and heterophily graphs, demonstrate the superiority of GCFormer in node classification, when compared to representative graph neural networks (GNNs) and graph Transformers. Jinsong Chen 0002, Hanpeng Liu, John E. Hopcroft, Kun He 0001 |
NeurIPS | 1 |
| 2024 | Neighborhood convolutional graph neural network
Jinsong Chen 0002, Boyu Li 0003, Kun He 0001 |
Knowl. Based Syst. | 1 |
| 2024 | PAMT: A Novel Propagation-Based Approach via Adaptive Similarity Mask for Node ClassificationabstractSemisupervised node classification on attributed networks is a crucial task for network analysis. By decoupling two critical operations in graph convolutional networks (GCNs), namely feature transformation and neighborhood aggregation, recent works of decoupled GCNs could support the information to propagate deeper and achieve advanced performance on node classification. However, they follow the structure-aware propagation strategy of GCNs, making it hard to capture the attribute correlation of nodes and be sensitive to the structure noise described by edges whose two endpoints belong to different categories. To address these issues, we propose a new method called the propagation with adaptive mask then training (PAMT). The key idea is to integrate the attribute similarity mask into the structure-aware propagation process. In this way, PAMT could preserve the attribute correlation of adjacent nodes during the propagation and effectively reduce the influence of structure noise. Moreover, we develop an iterative refinement mechanism to update the similarity mask during the training process to improve the training performance. Extensive experiments on six real-world datasets demonstrate the superior performance and robustness of PAMT over the state-of-the-art baselines. Jinsong Chen 0002, Boyu Li 0003, Qiuting He, Kun He 0001 |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2023 | NAGphormer: A Tokenized Graph Transformer for Node Classification in Large Graphs
Jinsong Chen 0002, Kaiyuan Gao, Gaichao Li, Kun He 0001 |
ICLR | 1 |
| 2022 | Structural Robust Label Propagation on Homogeneous GraphsabstractLabel propagation and graph neural networks are two main methods for the semi-supervised node classification problem on graphs. They share similar idea of propagating information over the network, exhibiting promising performance on the node classification task. Despite effectiveness, the limitations of these propagation methods are still not well understood. From the perspective of label propagation then training, we observe three major challenges of these propagation methods. The observations from both theoretical analyses and empirical studies reveal that the propagation operations can degrade performance on certain labels and suffers from structure noise, which is described by edges with two nodes belonging to distinct labels. To address the above issues, we propose a new method termed Robust Label Propagation (RLP). RLP contains two novel strategies, Robust Training and Ada-Mixup. Robust Training can utilize more attribute information to have the overall training correction ability and alleviate the impact of structure noise significantly. Ada-Mixup can help RLP mine useful structure information by integrating information before and after the propagation adaptively. Extensive empirical studies on real-world datasets demonstrate that RLP outperforms the mainstream baselines on the node classification task in terms of effectiveness, efficiency and robustness. Qiuting He, Jinsong Chen 0002, Hao Xu 0047, Kun He 0001 |
ICDM | 2 |