EDBT 2026 Demo / reviewers in the wild / expert
Jiaqing Fan
dblp:227/6586
· DBLP profile ↗
25ranked-venue papers
7as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | F2SST: Frequency-to-Spatial Semantic Transfer for Few-Shot Image ClassificationabstractFew-shot image classification (FSIC) aims to recognize novel categories from only a few labeled examples, making it inherently challenging under limited supervision. Existing approaches have attempted to alleviate this issue by incorporating explicit semantics like class names or knowledge graphs to guide learning. However, such methods often encounter semantic ambiguity due to their dependence on either overly simplistic semantic priors or resource-intensive external knowledge sources, which limits their potential. In this paper, we explore the frequency domain as an implicit and task-adaptive source of semantic information. We propose F2SST, a Frequency-to-Spatial Semantic Transfer framework that enhances feature learning by leveraging spectral signals as hidden semantics. Specifically, F2SST applies Fast Fourier Transform (FFT) to extract phase-invariant global frequency descriptors, followed by a lightweight Gated Spectral Attention (GSA) module that selectively emphasizes class-relevant frequency components. These enhanced spectral cues are then integrated into the spatial stream through a class-guided fusion mechanism, enabling more robust and semantically aligned representations. Extensive experiments on four standard benchmarks (miniImageNet, tieredImageNet, CIFAR-FS and FC100) demonstrate that F2SST consistently improves performance, validating the effectiveness of frequency-domain semantics in FSIC. Xueyi Chen, Bangjun Wang, Jiaqing Fan, Li Zhang 0004, Fanzhang Li |
AAAI | 3 |
| 2026 | Training-Free Spatio-temporal Decoupled Reasoning Video Segmentation with Adaptive Object MemoryabstractReasoning Video Object Segmentation (ReasonVOS) is a challenging task that requires stable object segmentation across video sequences using implicit and complex textual inputs. Previous methods fine-tune Multimodal Large Language Models (MLLMs) to produce segmentation outputs, which demand substantial resources. Additionally, some existing methods are coupled in the processing of spatio-temporal information, which affects the temporal stability of the model to some extent. To address these issues, we propose Training-Free Spatio-temporal Decoupled Reasoning Video Segmentation with Adaptive Object Memory (SDAM). We aim to design a training-free reasoning video segmentation framework that outperforms existing methods requiring fine-tuning, using only pre-trained models. Meanwhile, we propose an Adaptive Object Memory module that selects and memorizes key objects based on motion cues in different video sequences. Finally, we propose Spatio-temporal Decoupling for stable temporal propagation. In the spatial domain, we achieve precise localization and segmentation of target objects, while in the temporal domain, we leverage key object temporal information to drive stable cross-frame propagation. Our method achieves excellent results on five benchmark datasets, including Ref-YouTubeVOS, Ref-DAVIS17, MeViS, ReasonVOS, and ReVOS. Zhengtong Zhu, Jiaqing Fan, Zhixuan Liu, Fanzhang Li |
AAAI | 2 |
| 2026 | How Do Graph Signals Affect Recommendation: Unveiling the Mystery of Low and High-Frequency Graph SignalsabstractSpectral graph neural networks (GNNs) are highly effective in modeling graph signals, with their success in recommendation often attributed to low-pass filtering. However, recent studies highlight the importance of high-frequency signals. The role of low-frequency and high-frequency graph signals in recommendation remains unclear. This paper aims to bridge this gap by investigating the influence of graph signals on recommendation performance. We theoretically prove that the effects of low-frequency and high-frequency graph signals are equivalent in recommendation tasks, as both contribute by smoothing the similarities between user-item pairs. To leverage this insight, we propose a frequency signal scaler, a plug-and-play module that adjusts the graph signal filter function to fine-tune the smoothness between user-item pairs, making it compatible with any GNN model. Additionally, we identify and prove that graph embedding-based methods cannot fully capture the characteristics of graph signals. To address this limitation, a space flip method is introduced to restore the expressive power of graph embeddings. Remarkably, we demonstrate that either low-frequency or high-frequency graph signals alone are sufficient for effective recommendations. Extensive experiments on four public datasets validate the effectiveness of our proposed methods. Code is avaliable at https://github.com/mojosey/SimGCF. Feng Liu 0044, Hao Cang, Huanhuan Yuan, Jiaqing Fan, Yongjing Hao, Fuzhen Zhuang, Guanfeng Liu 0001, Pengpeng Zhao 0001 |
KDD (1) | 4 |
| 2026 | A local-global approach to point cloud semantic segmentation
RenJie Shen, Xiaoxia Xu 0004, Jiaqing Fan, Francisco Javier Cabrerizo |
Vis. Comput. | 3 |
| 2025 | Fuzzy Collaborative ReasoningabstractCollaborative reasoning enhances recommendation performance by combining the strengths of symbolic learning and deep neural learning. However, current collaborative reasoning models rely on parameterized networks to simulate logical operations within the reasoning process, which (1) do not comply with all axiomatic principles of classical logic and (2) limit the model's generalizability. To address these limitations, a Fuzzy logic approach tailored for Collaborative Reasoning (FuzzCR) is proposed in this work, aiming to augment the recommendation system with cognitive abilities. Specifically, this method redefines the sequential recommendation task as a logical query answering process to facilitate a more structured and logical progression of reasoning. Moreover, learning-free fuzzy logical operations are implemented for the designed reasoning process. Taking advantage of the inherent properties of fuzzy logic, these logical operations satisfy fundamental logical rules and ensure complete reasoning. After training, these operations can be applied to flexible reasoning processes, rather than being confined to fixed computation graphs, thereby exhibiting good generalizability. Extensive experiments conducted on publicly available datasets demonstrate the superiority of this method in solving the sequential recommendation task. Huanhuan Yuan, Pengpeng Zhao 0001, Jiaqing Fan, Junhua Fang, Guanfeng Liu 0001, Victor S. Sheng |
AAAI | 3 |
| 2025 | GPL4SRec: Graph Multi-Level Aware Prompt Learning for Streaming RecommendationabstractStreaming Recommendation (SRec) aims to capture evolving user preferences in the streaming scenarios. Recently, Graph Prompt Learning (GPL) methods have demonstrated their effectiveness and adaptability within SRec. However, existing graph prompt solutions rarely consider the evolution of multi-hop cascading relationships between users and items, which are crucial for modeling the shifts in user preferences. To address this problem, we propose a novel Graph Multi-Level Aware Prompt Learning for Streaming Recommendation, named GPL4SRec. Specifically, a graph encoder is first pre-trained on extensive historical data to capture user long-term preferences. Then, we design three types of prompts, namely node-aware, structure-aware, and layer-aware prompts, which are used to guide the pre-trained encoder to better capture user short-term preferences. This is accomplished by accounting for both the incremental changes in users and items, as well as the cascading evolution in multi-hop relationships. Furthermore, we provide a theoretical analysis showing that our prompt templates are critical to achieving superior performance. Finally, experimental results also prove that our model significantly outperforms the state-of-the-art approaches in SRec. Hao Cang, Huanhuan Yuan, Jiaqing Fan, Lei Zhao 0001, Guanfeng Liu 0001, Pengpeng Zhao 0001 |
IJCAI | 3 |
| 2025 | PeriodVOS: Learning Periodic Patterns for Unsupervised Video Object Segmentation via Adaptive Contextual Coupling
Jiaqing Fan, Hanwen Qian, Mengjuan Jiang, Fanzhang Li |
ACM Multimedia | 1 |
| 2025 | Graph Laplacian Regularized Referring Video Object Segmentation with Bayesian Neural Network Uncertainty Quantification
Jiaqing Fan, Fanzhang Li |
PRCV (11) | 2 |
| 2025 | Foundational and Specialized Continual Learning for Unsupervised Video Object Segmentation via Lie Group Structural Adapter
Hanwen Qian, Jiaqing Fan, Fanzhang Li |
PRCV (11) | 2 |
| 2025 | Enhancing few-shot class-incremental learning through prototype optimization
Mengjuan Jiang, Jiaqing Fan, Fanzhang Li |
Appl. Intell. | 2 |
| 2025 | Advances in continual learning: A comprehensive review
Mengjuan Jiang, Jiaqing Fan, Fanzhang Li |
Expert Syst. Appl. | 2 |
| 2025 | S2Trans: Structured spectrum transformer for robust unsupervised video object segmentation
Jiaqing Fan, Mengjuan Jiang, Fanzhang Li |
Neurocomputing | 2 |
| 2025 | Underwater Image Enhancement Based on U-Net Architecture and Channel Attention Mechanism Fusion Generative Adversarial NetworkabstractIn response to the challenges of blur distortion, low contrast and color fading in underwater images, caused by complex environmental factors and light attenuation, this study presents a novel underwater image enhancement method that leverages the U-Net architecture and channel attention mechanism fusion generative adversarial network (GAN), named UAEGAN. UAEGAN is built on the framework of GAN, combining the U-Net structure with a channel attention mechanism to construct a generator network, reducing the loss of low-level information during feature extraction and enhancing image details. Additionally, the algorithm employs a PatchGAN discriminator, which improves image resolution and detail representation by performing fine-grained true/false judgments on local image patches. Finally, the visual quality of the enhanced image is further optimized through the weighted fusion of multiple loss functions. Experimental results on the UIEB dataset indicate that UAEGAN outperforms the latest methods in terms of both visual quality and numerical metrics. The algorithm effectively enhances the clarity and visual quality of underwater images, providing strong support for subsequent underwater image processing tasks and applications. Jiaqing Fan, Chenyu Cheng |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2025 | Contrastive prototype network with prototype augmentation for few-shot classification
Mengjuan Jiang, Jiaqing Fan, Jiangzhen He, Weidong Du, Fanzhang Li |
Inf. Sci. | 2 |
| 2025 | Complementarily Learning Decoupled Category-Region-Aware Prototype for Few-Shot ClassificationabstractOpen-world few-shot classification is restricted by inadequate image-level content representation capabilities when the training and testing sets have significant differences in categories. Recently, many studies show the effectiveness of deep local descriptor-based methods, which attempt to select out dominating contents and discard noisy ones. However, aforementioned methods focus more on external relevance of support and query sets to filter features and ignore internal relevance among support sets, leading to unsatisfying classification performance. To relieve the issue, in this article, we propose the complementary learning Decoupling Category-Region-Aware Network (DCRNet) to simultaneously learn the correlation between internal members and then interact with the external sets. Specifically, we first propose an effective learnable Category Prototype-generated Feature Decoupling Module (CPFDM) to mine co-existing representations and generate comprehensive global class prototype. Then, to adaptively filter out discriminative local descriptors, we present a Category-Aware Selection Module (CASM) and introduce the Category-Aware Contrastive Loss (CACL) to highlight local information that is highly relative to the current category. In addition, the Region-Aware Contrastive Loss (RACL) is designed to encourage the model to concentrate on local regions, yielding powerful ability to distinguish foreground regions from between various categories. Finally, we leverage the filtered support descriptors to adaptively refine query descriptors through the descriptor selection strategy. Extensive experiments demonstrate that the proposed solution outperforms state-of-the-arts on five mainstream general and fine-grained few-shot classification datasets. We have released the training and testing code on https://github.com/jjfang007/DCRNet . Jiajie Fang, Mengjuan Jiang, Jiaqing Fan, Bangjun Wang, Fanzhang Li |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | DNTextSpotter: Arbitrary-Shaped Scene Text Spotting via Improved Denoising TrainingabstractMore and more end-to-end text spotting methods based on Transformer architecture have demonstrated superior performance. These methods utilize a bipartite graph matching algorithm to perform one-to-one optimal matching between predicted objects and actual objects. However, the instability of bipartite graph matching can lead to inconsistent optimization targets, thereby affecting the training performance of the model. Existing literature applies denoising training to solve the problem of bipartite graph matching instability in object detection tasks. Unfortunately, this denoising training method cannot be directly applied to text spotting tasks, as these tasks need to perform irregular shape detection tasks and more complex text recognition tasks than classification. To address this issue, we propose a novel denoising training method (DNTextSpotter) for arbitrary-shaped text spotting. Specifically, we decompose the queries of the denoising part into noised positional and noised content queries. We use the four Bezier control points of the Bezier center curve to generate the noised positional queries. For the noised content queries, considering that the output of the text in a fixed positional order is not conducive to aligning position with content, we employ a masked character sliding method to initialize noised content queries, thereby assisting in the alignment of text content and position. Additionally, to improve the model's perception of the background, we further utilize an additional loss function for background characters classification in the denoising training part. DNTextSpotter outperforms state-of-the-art methods on four benchmarks-Total-Text, SCUT-CTW1500, ICDAR15, and Inverse-Text-most notably achieving an 11.3% improvement over the best approach on Inverse-Text. Yu Xie 0001, Shaoyao Huang, Jiaqing Fan, Ziqiang Cao, Yue Zhang 0011 |
ACM Multimedia | 6 |
| 2024 | Dual temporal memory network with high-order spatio-temporal graph learning for video object segmentation
Jiaqing Fan, Shenglong Hu, Kaihua Zhang 0001, Bo Liu 0005 |
Image Vis. Comput. | 1 |
| 2023 | Temporally Efficient Gabor Transformer for Unsupervised Video Object SegmentationabstractSpatial-temporal structural details of targets in video (e.g. varying edges, textures over time) are essential to accurate Unsupervised Video Object Segmentation (UVOS). The vanilla multi-head self-attention in the Transformer-based UVOS methods usually concentrates on learning the general low-frequency information (e.g. illumination, color), while neglecting the high-frequency texture details, leading to unsatisfying segmentation results. To address this issue, this paper presents a Temporally efficient Gabor Transformer (TGFormer) for UVOS. The TGFormer jointly models the spatial dependencies and temporal coherence intra- and inter-frames, which can fully capture the rich structural details for accurate UVOS. Concretely, we first propose an effective learnable Gabor filtering Transformer to mine the structural texture details of the object for accurate UVOS. Then, to adaptively store the redundant neighboring historical information, we present an efficient dynamic neighboring frame selection module to automatically choose the useful temporal information, which simultaneously relieves the blurry frame and reduces the computation burden. Finally, we make the UVOS model be a fully Transformer architecture, meanwhile aggregating the information from space, Gabor and time domains, yielding a strong representation with rich structure details. Extensive experiments on five mainstream UVOS benchmarks (DAVIS2016, FBMS, DAVSOD, ViSal, and MCL) demonstrate the superiority of the presented solution to sate-of-the-art methods. Jiaqing Fan, Tiankang Su, Kaihua Zhang 0001, Bo Liu 0005, Qingshan Liu 0001 |
ACM Multimedia | 1 |
| 2022 | Bidirectionally Learning Dense Spatio-temporal Feature Propagation Network for Unsupervised Video Object SegmentationabstractSpatio-temporal feature representation is essential for accurate unsupervised video object segmentation, which needs an effective feature propagation paradigm for both appearance and motion features that can fully interchange information across frames. However, existing solutions mainly focus on the forward feature propagation from the preceding frame to the current one, either using the former segmentation mask or motion propagation in a frame-by-frame manner. This ignores the bi-directional temporal feature interactions (including the backward propagation from the future to the current frame) across all frames that can help to enhance the spatiotemporal feature representation for segmentation prediction. To this end, this paper presents a novel Dense Bidirectional Spatio-temporal feature propagation Network (DBSNet) to fully integrate the forward and the backward propagations across all frames. Specifically, a dense bi-ConvLSTM module is first developed to propagate the features across all frames in a forward and backward manner. This can fully capture the multi-level spatio-temporal contextual information across all frames, producing an effective feature representation that has a strong discriminative capability to tell from noisy backgrounds. Following it, a spatio-temporal Transformer refinement module is designed to further enhance the propagated features, which can effectively capture the spatio-temporal long-range dependencies among all frames. Afterwards, a Co-operative Direction-aware Graph Attention (Co-DGA) module is designed to integrate the propagated appearancemotion cues, yielding a strong spatio-temporal feature representation for segmentation mask prediction. The Co-DGA assigns proper attentional weights to neighboring points along the coordinate axis, making the segmentation model to selectively focus on the most relevant neighbors. Extensive evaluations on four mainstream challenging benchmarks including DAVIS16, FBMS, DAVSOD, and MCL demonstrate that the proposed DBSNet achieves favorable performance against state-of-the-art methods in terms of all evaluation metrics. Jiaqing Fan, Tiankang Su, Kaihua Zhang 0001, Qingshan Liu 0001 |
ACM Multimedia | 1 |
| 2022 | Semi-Supervised Video Object Segmentation via Learning Object-Aware Global-Local CorrespondenceabstractIn semi-supervised video object segmentation (VOS) task, temporal coherent object-level cues play a key role yet are hard to accurately model. To this end, this paper presents an object-aware global-local correspondence architecture, which enables to extract the inter-frame temporal coherent object-level features for accurate VOS. Specifically, we first generate a set of object masks by the ground-truth segmentation, and then we squeeze the current frame representation inside the object masks into a set of global object embeddings. Second, we compute the similarity between each embedding and the feature map, producing an object-aware weight for each pixel. The object-aware feature at each pixel is then constructed by summing the object embeddings weighted by their corresponding object-aware weights, which is able to capture rich object category information. Third, to establish the accurate correspondences between the inter-frame temporal coherent cues, we further design a novel global-local correspondence module to refine the temporal feature representations. Finally, we augment the object-aware features with the global-local aligned information to produce a strong spatio-temporal representation, which is essential to a more reliable pixel-wise segmentation prediction. Extensive evaluations are conducted on three popular VOS benchmarks containing Youtube-VOS, Davis2017 and Davis2016, demonstrating that the proposed method achieves favourable performance compared to the state-of-the-arts. Jiaqing Fan, Bo Liu 0005, Kaihua Zhang 0001, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Feature Alignment and Aggregation Siamese Networks for Fast Visual TrackingabstractSiamese networks have been successfully introduced into visual tracking, which match the best candidate and a target template via a couple of networks with shared parameters. However, most Siamese network-based trackers (SNTs) are tailored to best match the canonical posture of the template and the search-region images, resulting in inferior performance when the target objects have large-scale pose variations. Besides, SNTs fail to discriminate distractors well because they only leverage high-level semantic features as target representations that cannot well tell from different targets of the same category. To address these issues, this paper presents an efficient and effective SNT that is based on feature alignment and aggregation networks. Specifically, we first design an effective feature alignment network module to calibrate the search-region image. This module results in a more reliable matching response that is robust to severe target pose variations. Then, we develop an effective shallow-level and high-level feature aggregation network module to complement the feature characteristics, making the learned feature representation not only well differentiate the target from distractors, but also robust to target appearance variations. Afterwards, we employ a channel-attention mechanism to further strengthen the discriminative capability of the aggregated feature representation. Finally, both the alignment and the aggregation modules are seamlessly integrated into the Siamese networks for robust tracking. Meanwhile, we offline learn the network parameters end-to-end without time-consuming fine-tuning. Extensive evaluations on a variety of benchmarks including VOT-2017, OTB-100, UAV123 and GOT-10k demonstrate favorable performance of our tracker against state-of-the-art ones with a speed of 60 fps. Jiaqing Fan, Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Real-time manifold regularized context-aware correlation tracking
Jiaqing Fan, Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001, Wei Lian |
Frontiers Comput. Sci. | 1 |
| 2020 | Dynamically Spatiotemporal Regularized Correlation TrackingabstractRecently, due to the high performance, spatially regularized strategy has been widely applied to addressing the issue of boundary effects existed in correlation filter (CF)-based visual tracking. Specifically, it introduces a spatially regularized term to penalize the coefficients of the CFs to be learned depending on their spatial locations. However, the regularization weights are often formed as a fixed Gaussian function, and hence may cause the learned model degenerate due to the inflexible constraints on the ever-changing CFs to be learned over time during tracking. To address this issue, in this paper, we develop a dynamically spatiotemporal regularization model to constrain the CFs to be learned with the ever-changing regularization weights learned from two consecutive frames. The proposed method jointly learns the CFs along with the dynamically spatiotemporal constraint term, which can be efficiently solved in the Fourier domain by the alternative direction method. Extensive evaluations on the popular data sets OTB-100 and VOT-2016 demonstrate that the proposed tracker performs favorably against the baseline tracker and several recently proposed state-of-the-art methods. Yuhui Zheng, Huihui Song 0003, Kaihua Zhang 0001, Jiaqing Fan, Xinyan Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | Parallel Attentive Correlation TrackingabstractPsychological and cognitive findings indicate that human visual perception is attentive and selective, which may process spatial and appearance selective attentions in parallel. By reflecting some aspects of these attentions, this paper presents a novel correlation filter (CF) based tracking approach, corresponding to processing a local and a semi-local background domains, respectively. In the local domain, inspired by the Gestalt principle of figure-ground segregation, we leverage an efficient Boolean map representation, which characterizes an image by a set of Boolean maps via randomly thresholding its color channels, yielding a location response map as a weighted sum of all Boolean maps. The Boolean maps capture the topological structures of target and its scene with different granularities, thereby enabling to effectively improve tracking of non-rectangular objects. Alternatively, in the semi-local domains, we introduce a novel distractor-resilient metric regularization into CF, which acts as a force to push distractors into negative space. Consequently, the unwanted boundary effects of CF can be effectively alleviated. Finally, both models associated with the local and the semi-local domains are seamlessly integrated into a Bayesian framework, and the tracked location is determined by maximizing its likelihood function. Extensive evaluations on the OTB50, OTB100, VOT2016 and VOT2017 tracking benchmarks demonstrate that the proposed method achieves favorable performance against a variety of state-of-the-art trackers with a speed of 45 fps on a single CPU. Kaihua Zhang 0001, Jiaqing Fan, Qingshan Liu 0001, Jian Yang 0003, Wei Lian |
IEEE Trans. Image Process. | 2 |
| 2018 | Visual Tracking via Nonlocal Similarity LearningabstractEither global (e.g., intensity histograms and coefficients of sparse representation) or local (e.g., scale-invariant feature transform and histogram of oriented gradient) feature representations have been widely exploited for visual tracking. However, most of these representations describe a target appearance with a fixed spatial grid layout without considering the interactions between different grids, and hence may adversely affect their performance when the target appearance suffers from large-scale pose variations. In this paper, we learn a similarity function that considers the interactions of features in the grids not only from the same spatial positions, but also from different positions, thereby taking charge of the nonlocal information of the target appearances to effectively handle the significant appearance variations. Specifically, we explore the polynomial kernel feature map to characterize the nonlocal similarity information of all pairs of grids among the target and its background samples, and combine these feature maps as the target representations. Moveover, we learn a linear logistic regression classifier with online update to separate the target from its local background, and integrate this classifier into a particle filtering tracking framework. Extensive experimental results on the CVPR2013 tracking benchmark demonstrate the proposed approach performs favorably against some representative tracking algorithms. Qingshan Liu 0001, Jiaqing Fan, Huihui Song 0003, Kaihua Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |