Peijia Chen

dblp:244/9712 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2024
0009-0005-1743-2074ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Sparse Transformer Refinement Similarity Map for Aerial Tracking
abstract
Transformers have significantly enhanced tracking performance in visual tracking tasks. The key to its success lies in its powerful self-attention mechanism. Most current research focuses on using Transformers for feature extraction to concentrate on the target itself, with limited attention given to the refinement of similarity map. The refinement of the similarity map aims to adjust the focus on the target information itself. However, naive self-attention lacks the ability to prioritize the most important information and is easily influenced by other information. In this paper, we introduce a sparse attention mechanism to handle the refinement of the similarity map module, tackling this issue and achieving more accurate tracking. Additionally, we have constructed a more suitable encoder to effectively encode spatial features and temporal information. Extensive experiments demonstrate that our method (TCT+) achieves efficient aerial tracking, outperforming our baseline (TCTrack) on UAV123, DTB70, and UAV123@10fps. Our code is available at https://github.com/wolfwaytx/TCTPlus.
Ke Qi, Peijia Chen, Yutao Qi
ICIP3
2024 MFVG: A Visual Grounding Network with Multi-scale Fusion
abstract
Visual grounding, as a crucial multimodal reasoning task, aims to locate target objects in images based on natural language queries. This task requires the model to perform multimodal fusion and reasoning effectively. Early methods often rely on complex and manually designed modules for multimodal fusion and reasoning. However, these methods are usually customized for certain specific scenarios, thus limiting the generalization ability of the model. Recent works achieve visual grounding through the attention mechanism, which can capture the alignment relationship between vision and language, but ignore the importance of different scale features for multimodal reasoning. This paper proposes MFVG, a concise and effective visual grounding framework based on multiscale fusion guided by texts, which learns visual features with discriminative semantics through text queries. Specifically, MFVG allows the contextual semantic information of vision and language to interact fully and fuses features at different scales guided by text queries to capture richer detail features and semantic information, thereby enhancing the representational ability of the model and achieving better visual grounding. We conducted extensive experiments on five widely used benchmarks. The experiment results show that our proposed MFVG is superior to or comparable with the state-of-the-art methods.
Peijia Chen, Ke Qi, Jingdong Zhang 0002
ICMR1
2024 Multiple object tracking with segmentation and interactive multiple model
Ke Qi, Wenbin Chen 0003, Peijia Chen
J. Vis. Commun. Image Represent.5
2023 Activation to Saliency: Forming High-Quality Labels for Unsupervised Salient Object Detection
abstract
This paper focuses on the Unsupervised Salient Object Detection (USOD) issue. We come up with a two-stage Activation-to-Saliency (A2S) framework that effectively excavates saliency cues to train a robust saliency detector. It is worth noting that our method does not require any manual annotation in the whole process. In the first stage, we transform an unsupervisedly pre-trained network to aggregate multi-level features into a single activation map, where an Adaptive Decision Boundary (ADB) is proposed to assist the training of the transformed network. Moreover, a new loss function is proposed to facilitate the generation of high-quality pseudo labels. In the second stage, a self-rectification learning strategy is developed to train a saliency detector and refine the pseudo labels online. In addition, we construct a lightweight saliency detector using two Residual Attention Modules (RAMs) to learn robust saliency information. Extensive experiments on several SOD benchmarks prove that our framework reports significant performance compared with existing USOD methods. Moreover, training our framework on 3,000 images consumes about 1 hour, which is over 10 times faster than previous state-of-the-art methods. Code has been published athttps://github.com/moothes/A2S-USOD.
Huajun Zhou, Peijia Chen, Lingxiao Yang, Xiaohua Xie, Jian-Huang Lai
IEEE Trans. Circuits Syst. Video Technol.2
2021 Confidence-Guided Adaptive Gate and Dual Differential Enhancement for Video Salient Object Detection
abstract
Video salient object detection (VSOD) aims to locate and segment the most attractive object by exploiting both spatial cues and temporal cues hidden in video sequences. However, spatial and temporal cues are often unreliable in real-world scenarios, such as low-contrast foreground, fast motion, and multiple moving objects. To address these problems, we propose a new framework to adaptively capture available information from spatial and temporal cues, which contains Confidence-guided Adaptive Gate (CAG) modules and Dual Differential Enhancement (DDE) modules. For both RGB features and optical flow features, CAG estimates confidence scores supervised by the IoU between predictions and the ground truths to re-calibrate the information with a gate mechanism. DDE captures the differential feature representation to enrich the spatial and temporal information and generate the fused features. Experimental results on four widely used datasets demonstrate the effectiveness of the proposed method against thirteen state-of-the-art methods.
Peijia Chen, Jian-Huang Lai, Guangcong Wang, Huajun Zhou
ICME1
2021 Training Person Re-identification Networks with Transferred Images
Junkai Deng, Zhan-Xiang Feng, Peijia Chen, Jian-Huang Lai
PRCV (1)3