VLDB 2026 Research / reviewers in the wild / expert
Xinlong Chen
dblp:118/7338
· DBLP profile ↗
16ranked-venue papers
2as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Feature space variation-based active learning sample query strategy for graph deep learning
Xinlong Chen, Mingyu Lin, Yuzhuo Wang 0001, Jin Li 0032, Feiyang Ye 0002, Yanggeng Fu |
Expert Syst. Appl. | 1 |
| 2026 | Hybrid adaptive graph transformer: Integrating neighborhood diversity and structural distillation
Yanggeng Fu, Wen Meng, Jianzhi Zhuang, Xinlong Chen |
Knowl. Based Syst. | 4 |
| 2026 | UniTrain: A universal iterative semi-supervised training framework for graph representation learning
Xinlong Chen, Jin Li 0032, Yisong Huang 0002, Jianzhi Zhuang, Chenjunhao Shi, Zuhao Xu, Yanggeng Fu |
Neural Networks | 1 |
| 2026 | Curriculum-guided graph self-augmentation: A progressive deepening framework for GNNs
Qirong Zhang, Jin Li 0032, Xinlong Chen, Yanggeng Fu |
Neural Networks | 4 |
| 2026 | CurST-Net: Curriculum Learning Guided Spatial-Temporal Network for Traffic Flow PredictionabstractTraffic flow prediction is crucial to intelligent transportation systems (ITSs). However, the existing methods usually ignore the problem that the prediction difficulty of nodes in the traffic network is actually different. Besides, they fail to effectively handle the dilution of the original semantic information passed layer by layer when capturing the global dependency. To address these issues, this article proposes a curriculum learning guided spatial–temporal network (CurST-Net). Inspired by the human learning process, CurST-Net introduces a curriculum learning (CL) module that defines four metrics from multiple views to evaluate node difficulty and uses a training scheduler (TNS) to gradually introduce easy-to-difficult training nodes to the model to improve prediction ability. Moreover, we design a global spatial–temporal encoder that uses a multihead spatial–temporal attention mechanism and performs interlayer residual scaling on the original semantic information to efficiently capture the global spatial–temporal correlation of nodes. To the best of our knowledge, this is the first work that uses CL to solve the varying prediction difficulty of nodes in traffic flow prediction. The effectiveness of our model is validated through extensive experiments with 12 baseline models on three real-world traffic datasets. For the 1-h prediction horizon, the MAE values of our model on the three datasets are 15.19, 19.22, and 15.50, respectively. Our method yields an overall reduction of 6.51% in MAE across all three datasets compared to DSTAGNN, an advanced traffic flow prediction method. Additionally, the inference time on the PEMSD8 dataset is 5.44 s, requiring only 45% of the inference time of DSTAGNN, which demonstrates a significant improvement in inference speed. Shouming Chen, Xinlong Chen, Yisong Huang 0002, Yaru Su, Genggeng Liu, Yanggeng Fu |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2025 | Attention-guided Self-reflection for Zero-shot Hallucination Detection in Large Language ModelsabstractHallucination has emerged as a significant barrier to the effective application of Large Language Models (LLMs).In this work, we introduce a novel Attention-Guided SElf-Reflection (AGSER) approach for zero-shot hallucination detection in LLMs.The AGSER method utilizes attention contributions to categorize the input query into attentive and non-attentive queries.Each query is then processed separately through the LLMs, allowing us to compute consistency scores between the generated responses and the original answer.The difference between the two consistency scores serves as a hallucination estimator.In addition to its efficacy in detecting hallucinations, AGSER notably reduces computational overhead, requiring only three passes through the LLM and utilizing two sets of tokens.We have conducted extensive experiments with four widelyused LLMs across three different hallucination benchmarks, demonstrating that our approach significantly outperforms existing methods in zero-shot hallucination detection. Qiang Liu 0006, Xinlong Chen, Yue Ding 0009, Liang Wang 0001 |
EMNLP | 2 |
| 2025 | Mavors: Multi-granularity Video Representation for Multimodal Large Language ModelabstractLong-context video understanding in Multimodal Large Language Models (MLLMs) faces a critical challenge: balancing computational efficiency with the retention of fine-grained spatio-temporal patterns. Existing approaches (e.g., sparse sampling, dense sampling with low resolution, and token compression) suffer from significant information loss in temporal dynamics, spatial details, or subtle interactions, particularly in videos with complex motion or varying resolutions. To address this, we propose Mavors, a novel framework that introduces Multi-granularity video representation for holistic long-video modeling. Specifically, Mavors directly encodes raw video content into latent representations through two core components: 1) an Intra-chunk Vision Encoder (IVE) that preserves high-resolution spatial features via 3D convolutions and Vision Transformers, and 2) an Inter-chunk Feature Aggregator (IFA) that establishes temporal coherence across chunks using transformer-based dependency modeling with chunk-level rotary position encodings. Moreover, the framework unifies image and video understanding by treating images as single-frame videos via sub-image decomposition. Experiments across diverse benchmarks demonstrate Mavors' superiority in maintaining both spatial fidelity and temporal continuity, significantly outperforming existing methods in tasks requiring fine-grained spatio-temporal reasoning. Yang Shi 0009, Yushuo Guan, Yuanxing Zhang, Weihong Lin, Jingyun Hua, Xinlong Chen, Bohan Zeng, Wentao Zhang 0001, Wenjing Yang 0002, Di Zhang 0026 |
ACM Multimedia | 10 |
| 2025 | MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video ScenariosabstractMultimodal Large Language Models (MLLMs) have achieved considerable accuracy in Optical Character Recognition (OCR) from static images. However, their efficacy in video OCR is significantly diminished due to factors such as motion blur, temporal variations, and visual effects inherent in video content. To provide clearer guidance for training practical MLLMs, we introduce MME-VideoOCR benchmark, which encompasses a comprehensive range of video OCR application scenarios. MME-VideoOCR features 10 task categories comprising 25 individual tasks and spans 44 diverse scenarios. These tasks extend beyond text recognition to incorporate deeper comprehension and reasoning of textual content within videos. The benchmark consists of 1,464 videos with varying resolutions, aspect ratios, and durations, along with 2,000 meticulously curated, manually annotated question-answer pairs. We evaluate 18 state-of-the-art MLLMs on MME-VideoOCR, revealing that even the best-performing model (Gemini-2.5 Pro) achieves only an accuracy of 73.7%. Fine-grained analysis indicates that while existing MLLMs demonstrate strong performance on tasks where relevant texts are contained within a single or few frames, they exhibit limited capability in effectively handling tasks that demand holistic video comprehension. These limitations are especially evident in scenarios that require spatio-temporal reasoning, cross-frame information integration, or resistance to language prior bias. Our findings also highlight the importance of high-resolution visual input and sufficient temporal coverage for reliable OCR in dynamic video scenarios. Yang Shi 0009, Huanqian Wang, Wulin Xie, Huanyao Zhang, Lijie Zhao, Yifan Zhang 0004, Xinfeng Li, Chaoyou Fu, Zhuoer Wen, Zhuoran Zhang 0003, Xinlong Chen, Bohan Zeng, Yushuo Guan, Zhang Zhang 0001, Liang Wang 0001, Haoxuan Li 0001, Zhouchen Lin, Yuanxing Zhang, Pengfei Wan 0001, Haotian Wang 0001, Wenjing Yang 0002 |
NeurIPS | 12 |
| 2025 | A Novel YJQR-LSTM Model for Nonparametric Probabilistic Sustainable Agriculture Wind Power Forecasting Based on Intelligent IoTabstractEnergy costs associated with the consumption of nonrenewable energy sources have become an important issue in improving the international competitiveness of agriculture. Wind power, as a renewable energy source, can replace nonrenewable energy sources to reduce energy costs and improve the sustainability of agricultural. However, the inherent intermittency, randomness, and volatility within weather conditions and wind speed present a substantial challenge in accurately predicting wind power generation. This work proposes a novel YJQR-LSTM algorithm that leverages Yeo-Johnson quantile regression (YJQR) with a long short-term memory (LSTM) network for nonparametric probabilistic forecasting of wind power generation via the Intelligent Internet of Things. First, an improved YJQR model based on the YJ transformation is designed to obtain a more precise characterization of wind uncertainty, providing a more flexible probability density function for wind power generation. Then, utilizing the unique structure of the LSTM network to learn the parameters of the YJQR model, temporal features can be extracted from time-series data. To mitigate the impact of outliers in the raw data on accuracy and improve computational efficiency, a novel logarithmic-likelihood function is developed as the loss function utilized in the training phase. The effectiveness of the proposed algorithm is validated using a real-world dataset from five wind farms from the Global Energy Forecasting Competition. Numerical results demonstrate that the algorithm provides more accurate wind power prediction results in complex wind power data environments, which is important for making full use of wind energy and thus reducing the consumption of nonrenewable energy in agriculture. Jie Wang 0163, Junhui Jiang 0001, Xinlong Chen, Defu Cai, Yue Wu 0026, Renzhi Lu |
IEEE Internet Things J. | 3 |
| 2025 | SE-GSSL: Soft-Mask enhanced graph self-supervised learning with multi-aspect knowledge encoding and adaptive sample selection
Yanggeng Fu, Xinlong Chen, Shuling Xu, Qirong Zhang, Wen Meng, Genggeng Liu |
Knowl. Based Syst. | 2 |
| 2025 | GSSCL: A framework for Graph Self-Supervised Curriculum Learning based on clustering label smoothing
Yanggeng Fu, Xinlong Chen, Shuling Xu, Jin Li 0032 |
Neural Networks | 2 |
| 2024 | Curriculum-Enhanced Residual Soft An-Isotropic Normalization for Over-Smoothness in Deep GNNsabstractDespite Graph neural networks' significant performance gain over many classic techniques in various graph-related downstream tasks, their successes are restricted in shallow models due to over-smoothness and the difficulties of optimizations among many other issues. In this paper, to alleviate the over-smoothing issue, we propose a soft graph normalization method to preserve the diversities of node embeddings and prevent indiscrimination due to possible over-closeness. Combined with residual connections, we analyze the reason why the method can effectively capture the knowledge in both input graph structures and node features even with deep networks. Additionally, inspired by Curriculum Learning that learns easy examples before the hard ones, we propose a novel label-smoothing-based learning framework to enhance the optimization of deep GNNs, which iteratively smooths labels in an auxiliary graph and constructs many gradual non-smooth tasks for extracting increasingly complex knowledge and gradually discriminating nodes from coarse to fine. The method arguably reduces the risk of overfitting and generalizes better results. Finally, extensive experiments are carried out to demonstrate the effectiveness and potential of the proposed model and learning framework through comparison with twelve existing baselines including the state-of-the-art methods on twelve real-world node classification benchmarks. Jin Li 0032, Qirong Zhang, Shuling Xu, Xinlong Chen, Longkun Guo, Yanggeng Fu |
AAAI | 4 |
| 2024 | Training Graph Transformers via Curriculum-Enhanced Attention DistillationabstractRecent studies have shown that Graph Transformers (GTs) can be effective for specific graph-level tasks. However, when it comes to node classification, training GTs remains challenging, especially in semi-supervised settings with a severe scarcity of labeled data. Our paper aims to address this research gap by focusing on semi-supervised node classification. To accomplish this, we develop a curriculum-enhanced attention distillation method that involves utilizing a Local GT teacher and a Global GT student. Additionally, we introduce the concepts of in-class and out-of-class and then propose two improvements, out-of-class entropy and top-k pruning, to facilitate the student's out-of-class exploration under the teacher's in-class guidance. Taking inspiration from human learning, our method involves a curriculum mechanism for distillation that initially provides strict guidance to the student and gradually allows for more out-of-class exploration by a dynamic balance. Extensive experiments show that our method outperforms many state-of-the-art approaches on seven public graph benchmarks, proving its effectiveness. Yisong Huang 0002, Jin Li 0032, Xinlong Chen, Yanggeng Fu |
ICLR | 3 |
| 2024 | LightCapsGNN: light capsule graph neural network for graph classification
Yucheng Yan, Shuling Xu, Xinlong Chen, Genggeng Liu, Yanggeng Fu |
Knowl. Inf. Syst. | 4 |
| 2024 | DWSSA: Alleviating over-smoothness for deep Graph Neural Networks
Qirong Zhang, Jin Li 0032, Qingqing Ye 0001, Yuxi Lin, Xinlong Chen, Yanggeng Fu |
Neural Networks | 5 |
| 2023 | Graph Contrastive Representation Learning with Input-Aware and Cluster-Aware Regularization
Jin Li 0032, Bingshi Li, Qirong Zhang, Xinlong Chen, Longkun Guo, Yanggeng Fu |
ECML/PKDD (2) | 4 |