Jiapei Fan

dblp:189/7713 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0003-3194-025XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Learning from Graph: Mitigating Label Noise on Graph through Topological Feature Reconstruction
abstract
Graph Neural Networks (GNNs) have shown remarkable performance in modeling graph data. However, Labeling graph data typically relies on unreliable information, leading to noisy node labels. Existing approaches for GNNs under Label Noise (GLN) employ supervision signals beyond noisy labels for robust learning. While empirically effective, they tend to over-reliance on supervision signals built upon external assumptions, leading to restricted applicability. In this work, we shift the focus to exploring how to extract useful information and learn from the graph itself, thus achieving robust graph learning. From an information theory perspective, we theoretically and empirically demonstrate that the graph itself contains reliable information for graph learning under label noise. Based on these insights, we propose the Topological Feature Reconstruction (TFR) method. Specifically, TFR leverages the fact that the pattern of clean labels can more accurately reconstruct graph features through topology, while noisy labels cannot. TFR is a simple and theoretically guaranteed model for robust graph learning under label noise. We conduct extensive experiments across datasets with varying properties. The results demonstrate the robustness and broad applicability of our proposed TFR compared to state-of-the-art baselines. Codes are available at https://github.com/eaglelab-zju/TFR.
Zhonghao Wang 0002, Yuanchen Bei, Sheng Zhou 0004, Zhiyao Zhou, Jiapei Fan, Hui Xue 0001, Haishuai Wang, Jiajun Bu
CIKM5
2025 Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction Tuning
abstract
Continual multimodal instruction tuning is crucial for adapting Multimodal Large Language Models (MLLMs) to evolving tasks. However, most existing methods adopt a fixed architecture, struggling with adapting to new tasks due to static model capacity. We propose to evolve the architecture under parameter budgets for dynamic task adaptation, which remains unexplored and imposes two challenges: 1) task architecture conflict, where different tasks require varying layer-wise adaptations, and 2) modality imbalance, where different tasks rely unevenly on modalities, leading to unbalanced updates. To address these challenges, we propose a novel Dynamic Mixture of Curriculum LoRA Experts (D-MoLE) method, which automatically evolves MLLM's architecture with controlled parameter budgets to continually adapt to new tasks while retaining previously learned knowledge. Specifically, we propose a dynamic layer-wise expert allocator, which automatically allocates LoRA experts across layers to resolve architecture conflicts, and routes instructions layer-wisely to facilitate knowledge sharing among experts. Then, we propose a gradient-based inter-modal continual curriculum, which adjusts the update ratio of each module in MLLM based on the difficulty of each modality within the task to alleviate the modality imbalance problem. Extensive experiments show that D-MoLE significantly outperforms state-of-the-art baselines, achieving a 15 percent average improvement over the best baseline. To the best of our knowledge, this is the first study of continual learning for MLLMs from an architectural perspective.
Chendi Ge, Xin Wang 0019, Zeyang Zhang 0001, Hong Chen 0011, Jiapei Fan, Longtao Huang, Hui Xue 0001, Wenwu Zhu 0001
ICML5
2025 Correlation-Aware Graph Convolutional Networks for Multi-Label Node Classification
abstract
Multi-label node classification is an important yet under-explored domain in graph mining as many real-world nodes belong to multiple categories rather than just a single one. Although a few efforts have been made by utilizing Graph Convolution Networks (GCNs) to learn node representations and model correlations between multiple labels in the embedding space, they still suffer from the ambiguous feature and ambiguous topology induced by multiple labels, which reduces the credibility of the messages delivered in graphs and overlooks the label correlations on graph data. Therefore, it is crucial to reduce the ambiguity and empower the GCNs for accurate classification. However, this is quite challenging due to the requirement of retaining the distinctiveness of each label while fully harnessing the correlation between labels simultaneously. To address these issues, in this paper, we propose a Correlation-aware Graph Convolutional Network (CorGCN) for multi-label node classification. By introducing a novel Correlation-Aware Graph Decomposition module, CorGCN can learn a graph that contains rich label-correlated information for each label. It then employs a Correlation-Enhanced Graph Convolution to model the relationships between labels during message passing to further bolster the classification process. Extensive experiments on five datasets demonstrate the effectiveness of our proposed CorGCN.
Yuanchen Bei, Weizhi Chen, Hao Chen 0062, Sheng Zhou 0004, Carl Yang 0001, Jiapei Fan, Longtao Huang, Jiajun Bu
KDD (1)6
2024 Large Language Model with Curriculum Reasoning for Visual Concept Recognition
abstract
Visual concept recognition aims to capture the basic attributes of an image and reason about the relationships among them to determine whether the image satisfies a certain concept, and has been widely used in various tasks such as human action recognition and image risk warning. Most existing works adopt deep neural networks for visual concept recognition, which are black-box and incomprehensible to humans, thus making them unacceptable for sensitive domains such as prohibited event detection and risk early warning etc. To address this issue, we propose to combine large language model (LLM) with explainable symbolic reasoning via curriculum reweighting to increase the interpretability and accuracy of visual concept recognition in this paper. However, realizing this goal is challenging given that i) the performance of symbolic representations are limited by the lack of annotated reasoning symbols and rules for most tasks, and ii) the LLMs may suffer from knowlege hallucination and dynamic open environment. To address these issues, in this paper, we propose CurLLM-Reasoner, a curriculum reasoning method based on symbolic reasoning and large language model for visual concept recognition. Specifically, we propose a novel rule enhancement module with a tool library, which fully leverage the reasoning capability of large language models and can generate human-understandable rules without any annotation. We further propose a curriculum data resampling methodology to help the large language model accurately extract from easy to complex rules at different reasoning stages. Extensive experiments on various datasets demonstrate that CurLLM-Reasoner can achieve the state-of-the-art visual concept recognition results with explainable rules while free of human annotations.
Yipeng Zhang 0003, Xin Wang 0019, Hong Chen 0011, Jiapei Fan, Weigao Wen, Hui Xue 0001, Hong Mei 0001, Wenwu Zhu 0001
KDD4
2024 NoisyGL: A Comprehensive Benchmark for Graph Neural Networks under Label Noise
abstract
Graph Neural Networks (GNNs) exhibit strong potential in node classification task through a message-passing mechanism. However, their performance often hinges on high-quality node labels, which are challenging to obtain in real-world scenarios due to unreliable sources or adversarial attacks. Consequently, label noise is common in real-world graph data, negatively impacting GNNs by propagating incorrect information during training. To address this issue, the study of Graph Neural Networks under Label Noise (GLN) has recently gained traction. However, due to variations in dataset selection, data splitting, and preprocessing techniques, the community currently lacks a comprehensive benchmark, which impedes deeper understanding and further development of GLN. To fill this gap, we introduce NoisyGL in this paper, the first comprehensive benchmark for graph neural networks under label noise. NoisyGL enables fair comparisons and detailed analyses of GLN methods on noisy labeled graph data across various datasets, with unified experimental settings and interface. Our benchmark has uncovered several important insights that were missed in previous research, and we believe these findings will be highly beneficial for future studies. We hope our open-source benchmark library will foster further advancements in this field. The code of the benchmark can be found in https://github.com/eaglelab-zju/NoisyGL.
Zhonghao Wang 0002, Danyu Sun, Sheng Zhou 0004, Haobo Wang 0001, Jiapei Fan, Longtao Huang, Jiajun Bu
NeurIPS5
2023 Continual Few-shot Learning with Transformer Adaptation and Knowledge Regularization
abstract
Continual few-shot learning, as a paradigm that simultaneously solves continual learning and few-shot learning, has become a challenging problem in machine learning. An eligible continual few-shot learning model is expected to distinguish all seen classes upon new categories arriving, where each category only includes very few labeled data. However, existing continual few-shot learning methods only consider the visual modality, where the distributions of new categories often indistinguishably overlap with old categories, thus resulting in the severe catastrophic forgetting problem. To tackle this problem, in this paper we study continual few-shot learning with the assistance of semantic knowledge by simultaneously taking both visual modality and semantic concepts of categories into account. We propose a Continual few-shot learning algorithm with Semantic knowledge Regularization (CoSR) for adapting to the distribution changes of visual prototypes through a Transformer-based prototype adaptation mechanism. Specifically, the original visual prototypes from the backbone are fed into the well-designed Transformer with corresponding semantic concepts, where the semantic concepts are extracted from all categories. The semantic-level regularization forces the categories with similar semantics to be closely distributed, while the opposite ones are constrained to be far away from each other. The semantic regularization improves the model’s ability to distinguish between new and old categories, thus significantly mitigating the catastrophic forgetting problem in continual few-shot learning. Extensive experiments on CIFAR100, miniImageNet, CUB200 and an industrial dataset with long-tail distribution demonstrate the advantages of our CoSR model compared with state-of-the-art methods.
Xin Wang 0019, Yue Liu 0025, Jiapei Fan, Weigao Wen, Hui Xue 0001, Wenwu Zhu 0001
WWW3
2017 Spatiotemporal correlation in WebGIS group-user intensive access patterns
abstract
Group-user intensive access to WebGIS exhibits spatiotemporal behaviour patterns with aggregation features and regularity distributions when geospatial data are accessed repeatedly over time and aggregated in certain spatial areas. We argue that these observable group-user access patterns provide a foundation for improved optimization of WebGIS so that it can respond to volume intensive requests with a higher quality of service and improve performance. Subsequently, a measure of access popularity distribution must precisely reflect the access aggregation and regularity features found in group-user intensive access. In our research, we considered both the temporal distribution characteristics and spatial correlation in the access popularity of tiled geospatial data (tiles). Based on the observation that group-user access follows a Zipf-like law, we built a tile-access popularity distribution based on time-sequence, to express the access aggregation of group-users with heavy-tailed characteristics. Considering the spatial locality of user-browsed tiles, we built a quantitative expression for the correlation between tile-access popularities and the distances to hotspot tiles, reflecting the attenuation of tile-access popularity to distance. Moreover, given the geographical spatial dependency and scale attribute of tiles, and the time-sequence of tile-access popularity, we built a Poisson regression model to express the degree of correlation among the accesses to adjacent tiles at different scales, reflecting the spatiotemporal correlation in tile access patterns. Experiments verify the accuracy of our Poisson regression model, which we then applied to a cluster-based cache-prefetching scenario. The results show that our model successfully reflects the spatiotemporal aggregation features of group-user intensive access and group-user behaviour patterns in WebGIS. The refined mathematical method in our model represents a time-sequence distribution of intensive access to tiles and the spatial aggregation and correlation in access to tiles at different scales, quantitatively expressing group-user spatiotemporal behaviour patterns with aggregation features and a regular distribution. Our proposed model provides a precise and empirical basis for performance-optimization strategies in WebGIS services, such as planning computing resource allocation and utilization, distributed storage of geospatial data, and providing distributed services so as to respond rapidly to geospatial data requests, thus addressing the challenges of volume-intensive user access.
Rui Li 0046, Jiapei Fan, Jie Jiang 0014, Huayi Wu
Int. J. Geogr. Inf. Sci.2