Jinqing Qi

dblp:09/287 · DBLP profile ↗
← Back
28ranked-venue papers
1as first author
15since 2021 · last 2025
0000-0002-3777-2405ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 TrackFusion: Enhancing Multi-Object Tracking With Temporal Trajectory Modeling and Frame-Integrated Detection
abstract
Although MOTIP is the SOTA multi-object tracking method, there are still some issues that limit its performance. First, MOTIP still has defects in temporal information modeling, which leads to the failure to fully utilize the historical information of the tracked target and affects the correlation performance of the model. Second, in MOT, objects in consecutive video frames usually have temporal continuity and spatial consistency. Therefore, the object information of the previous frame can effectively assist the detection of the current frame. However, MOTIP performs independent detection between each frame, which does not fully utilize the correlation information between frames, resulting in suboptimal model performance. To address the above problems, we propose TrackFusion, which optimizes model performance from the perspective of trajectory modeling and inter-frame joint detection. First, we extract embeddings in video sequences through a Transformer-based detector, then combine the embeddings of the same object in different frames into sequences and input them into the trajectory modeling module for sequence association. This strategy effectively enhances the association ability. Thanks to these improvements, TrackFusion’s HOTA on the DanceTrack test set reached 68.6%, an increase of 1.1% compared to MOTIP’s 67.5%.
Shuai Liu 0009, Bingyang Wang, Jiaojiao Dai, Jinqing Qi, Huchuan Lu, You He 0002
ICASSP5
2025 Towards Survivability in Complex Motion Scenarios: RGB-Event Object Tracking via Historical Trajectory Prompting
abstract
Event data has recently emerged as a valuable complement to object tracking, offering dense temporal resolution and a high dynamic range. However, existing RGB-Event trackers struggle with targets exhibiting complex motion trajectories, where RGB features alone fail to provide sufficient discrimination. To address this, we propose EventTPT, an innovative RGB-Event tracking framework that leverages pivotal prompts embedded in historical trajectories for enhanced tracking. Specifically, EventTPT integrates the trajectories of multiple adjacent frames into a single event image using a time-weighted aggregation and subsequently inputs this as a visual prompt into the tracker for current frame locating. A cross-modal adaptive fusion module is further designed for object perception in scenarios with photometric inconsistency. Additionally, we introduce EventUAV, a novel and challenging RGB-Event tracking benchmark featuring objects with intricate motion dynamics and poor visibility in RGB-only modalities. Extensive experiments demonstrate that EventTPT surpasses state-of-the-art trackers on EventUAV and achieves competitive performance on other benchmarks (e.g., COESOT and VisEvent), underscoring its strong generalizability and robustness for resilient robotic vision systems. The code can be found at https://github.com/xiawenhao2022/EventTPT.
Wenhao Xia, Jiawen Zhu 0003, Jinqing Qi, You He 0002, Xu Jia 0012
ICRA4
2025 Progressive Query-Driven Learning for Few-Shot Semantic Segmentation in Remote Sensing
Jinqing Qi
PRCV (15)2
2025 Learning Motion and Temporal Cues for Unsupervised Video Object Segmentation
abstract
In this article, we address the challenges in unsupervised video object segmentation (UVOS) by proposing an efficient algorithm, termed MTNet, which concurrently exploits motion and temporal cues. Unlike previous methods that focus solely on integrating appearance with motion or on modeling temporal relations, our method combines both aspects by integrating them within a unified framework. MTNet is devised by effectively merging appearance and motion features during the feature extraction process within encoders, promoting a more complementary representation. To capture the intricate long-range contextual dynamics and information embedded within videos, a temporal transformer module is introduced, facilitating efficacious interframe interactions throughout a video clip. Furthermore, we employ a cascade of decoders all feature levels across all feature levels to optimally exploit the derived features, aiming to generate increasingly precise segmentation masks. As a result, MTNet provides a strong and compact framework that explores both temporal and cross-modality knowledge to robustly localize and track the primary object accurately in various challenging scenarios efficiently. Extensive experiments across diverse benchmarks conclusively show that our method not only attains state-of-the-art performance in UVOS but also delivers competitive results in video salient object detection (VSOD). These findings highlight the method's robust versatility and its adeptness in adapting to a range of segmentation tasks. The source code is available at https://github.com/hy0523/MTNet.
Yunzhi Zhuge, Hongyu Gu, Lu Zhang 0053, Jinqing Qi, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.4
2024 Hybrid-SORT: Weak Cues Matter for Online Multi-Object Tracking
abstract
Multi-Object Tracking (MOT) aims to detect and associate all desired objects across frames. Most methods accomplish the task by explicitly or implicitly leveraging strong cues (i.e., spatial and appearance information), which exhibit powerful instance-level discrimination. However, when object occlusion and clustering occur, spatial and appearance information will become ambiguous simultaneously due to the high overlap among objects. In this paper, we demonstrate this long-standing challenge in MOT can be efficiently and effectively resolved by incorporating weak cues to compensate for strong cues. Along with velocity direction, we introduce the confidence and height state as potential weak cues. With superior performance, our method still maintains Simple, Online and Real-Time (SORT) characteristics. Also, our method shows strong generalization for diverse trackers and scenarios in a plug-and-play and training-free manner. Significant and consistent improvements are observed when applying our method to 5 different representative trackers. Further, with both strong and weak cues, our method Hybrid-SORT achieves superior performance on diverse benchmarks, including MOT17, MOT20, and especially DanceTrack where interaction and severe occlusion frequently happen with complex motions. The code and models are available at https://github.com/ymzis69/HybridSORT.
Mingzhan Yang, Guangxin Han, Bin Yan 0004, Jinqing Qi, Huchuan Lu, Dong Wang 0004
AAAI5
2023 Few-shot Semantic Segmentation by Exploiting Dynamic and Regional Contexts
abstract
Few-shot Semantic Segmentation (FSS) has received increasing interests recently. Modeling effective interaction be-tween support and query images is a crucial challenge in existing prototype based methods. In this paper, we propose a Dynamic and Regional Context Network (DRCNet) to achieve sufficient support-query interaction for accurate FSS. A Dynamic Context Module (DCM) is first proposed to capture the spatial details in query images by building dynamic convolutions in local views. To further alleviate the undesirable noises, a Regional Context Module (RCM) is proposed to mine and exclude the background and ambiguous objects in query images by modeling the prototypes for ambiguous regions. Experimental results on Pascal-5iand COCO-20idatasets demonstrate that our proposed DRCNet performs significantly superior against state-of-the-art methods.
Hongyu Gu, Yunzhi Zhuge, Lu Zhang 0053, Jinqing Qi, Huchuan Lu
ICME4
2022 You Only Infer Once: Cross-Modal Meta-Transfer for Referring Video Object Segmentation
abstract
We present YOFO (You Only inFer Once), a new paradigm for referring video object segmentation (RVOS) that operates in an one-stage manner. Our key insight is that the language descriptor should serve as target-specific guidance to identify the target object, while a direct feature fusion of image and language can increase feature complexity and thus may be sub-optimal for RVOS. To this end, we propose a meta-transfer module, which is trained in a learning-to-learn fashion and aims to transfer the target-specific information from the language domain to the image domain, while discarding the uncorrelated complex variations of language description. To bridge the gap between the image and language domains, we develop a multi-scale cross-modal feature mining block that aggregates all the essential features required by RVOS from both domains and generates regression labels for the meta-transfer module. The whole system can be trained in an end-to-end manner and shows competitive performance against state-of-the-art two-stage approaches.
Dezhuang Li, Ruoqi Li, Lijun Wang 0001, Yifan Wang 0004, Jinqing Qi, Lu Zhang 0053, Ting Liu 0018, Qingquan Xu, Huchuan Lu
AAAI5
2022 Depth-inspired Label Mining for Unsupervised RGB-D Salient Object Detection
abstract
Existing deep learning-based unsupervised Salient Object Detection (SOD) methods heavily rely on the pseudo labels predicted from handcrafted features. However, the pseudo ground truth obtained only from RGB space would easily bring undesirable noises, especially in some complex scenarios. This naturally leads to the incorporation of extra depth modality with RGB images for more robust object identification, namely RGB-D SOD. Compared with the well-studied unsupervised SOD in the RGB domain, deep unsupervised RGB-D SOD is a less explored direction in the literature. In this paper, we propose to tackle this task by introducing a novel systemic design for high-quality pseudo-label mining. Our framework consists of two key components, Depth-inspired Label Generation (DLG) and Multi-source Uncertainty-aware Label Optimization (MULO). In DLG, a lightweight deep network is designed for automatically producing pseudo labels from depth maps in a self-supervised manner. Then, MULO introduces an effective pseudo label optimization strategy by learning the uncertainty of the pseudo labels from the depth domain and heuristic features. Extensive experiments demonstrate that the proposed method significantly outperforms the state-of-the-art unsupervised methods on mainstream benchmarks.
Yue Wang 0038, Lu Zhang 0053, Jinqing Qi, Huchuan Lu
ACM Multimedia4
2022 Few-Shot Segmentation via Rich Prototype Generation and Recurrent Prediction Enhancement
Hongsheng Wang, Xiaoqi Zhao 0003, Youwei Pang, Jinqing Qi
PRCV (4)4
2022 Road extraction from satellite images with iterative cross-task feature enhancement
Weiling Yin, Mingyang Qian, Lijun Wang 0001, Jinqing Qi, Huchuan Lu
Neurocomputing4
2022 Learning to Detect Salient Object With Multi-Source Weak Supervision
abstract
High-cost pixel-level annotations makes it appealing to train saliency detection models with weak supervision. However, a single weak supervision source hardly contain enough information to train a well-performing model. To this end, we introduce a unified two-stage framework to learn from category labels, captions, web images and unlabeled images. In the first stage, we design a classification network (CNet) and a caption generation network (PNet), which learn to predict object categories and generate captions, respectively, meanwhile highlights the potential foreground regions. We present an attention transfer loss to transmit supervisions between two tasks and an attention coherence loss to encourage the networks to detect generally salient regions instead of task-specific regions. In the second stage, we create two complementary training datasets using CNet and PNet, i.e., natural image dataset with noisy labels for adapting saliency prediction network (SNet) to natural image input, and synthesized image dataset by pasting objects on background images for providing SNet with accurate ground-truth. During the testing phases, we only need SNet to predict saliency maps. Experiments indicate the performance of our method compares favorably against unsupervised, weakly supervised methods and even some supervised methods.
Hongshuang Zhang, Yu Zeng 0001, Huchuan Lu, Lihe Zhang, Jinqing Qi
IEEE Trans. Pattern Anal. Mach. Intell.6
2021 Learning Motion-Appearance Co-Attention for Zero-Shot Video Object Segmentation
abstract
How to make the appearance and motion information interact effectively to accommodate complex scenarios is a fundamental issue in flow-based zero-shot video object segmentation. In this paper, we propose an Attentive Multi-Modality Collaboration Network (AMC-Net) to utilize appearance and motion information uniformly. Specifically, AMC-Net fuses robust information from multi-modality features and promotes their collaboration in two stages. First, we propose a Multi-Modality Co-Attention Gate (MCG) on the bilateral encoder branches, in which a gate function is used to formulate co-attention scores for balancing the contributions of multi-modality features and suppressing the redundant and misleading information. Then, we propose a Motion Correction Module (MCM) with a visualmotion attention mechanism, which is constructed to emphasize the features of foreground objects by incorporating the spatio-temporal correspondence between appearance and motion cues. Extensive experiments on three public challenging benchmark datasets verify that our proposed network performs favorably against existing state-of-the-art methods via training with fewer data. The code is released at https://github.com/isyangshu/AMC-Net.
Shu Yang 0004, Lu Zhang 0053, Jinqing Qi, Huchuan Lu
ICCV3
2021 HAT: Hierarchical Aggregation Transformers for Person Re-identification
abstract
Recently, with the advance of deep Convolutional Neural Networks (CNNs), person Re-Identification (Re-ID) has witnessed great success in various applications.However, with limited receptive fields of CNNs, it is still challenging to extract discriminative representations in a global view for persons under non-overlapped cameras.Meanwhile, Transformers demonstrate strong abilities of modeling long-range dependencies for spatial and sequential data.In this work, we take advantages of both CNNs and Transformers, and propose a novel learning framework named Hierarchical Aggregation Transformer (HAT) for image-based person Re-ID with high performance.To achieve this goal, we first propose a Deeply Supervised Aggregation (DSA) to recurrently aggregate hierarchical features from CNN backbones.With multi-granularity supervision, the DSA can enhance multi-scale features for person retrieval, which is very different from previous methods.Then, we introduce a Transformer-based Feature Calibration (TFC) to integrate low-level detail information as the global prior for high-level semantic information.The proposed TFC is inserted to each level of hierarchical features, resulting in great performance improvements.To our best knowledge, this work is the first to take advantages of both CNNs and Transformers for image-based person Re-ID.Comprehensive experiments on four large-scale Re-ID benchmarks demonstrate that our method shows better results than several state-of-the-art methods.The code is released at https://github.com/AI-Zhpp/HAT.
Guowen Zhang, Jinqing Qi, Huchuan Lu
ACM Multimedia3
2021 Learning Regression and Verification Networks for Robust Long-term Tracking
Lijun Wang 0001, Dong Wang 0004, Jinqing Qi, Huchuan Lu
Int. J. Comput. Vis.4
2021 Self-attention guided representation learning for image-text matching
Xuefei Qi, Ying Zhang 0021, Jinqing Qi, Huchuan Lu
Neurocomputing3
2020 Multi-attention guided feature fusion network for salient object detection
Anni Li, Jinqing Qi, Huchuan Lu
Neurocomputing2
2019 Language-aware weak supervision for salient object detection
Mingyang Qian, Jinqing Qi, Lihe Zhang, Mengyang Feng, Huchuan Lu
Pattern Recognit.2
2019 Edge-Aware Convolution Neural Network Based Salient Object Detection
abstract
Salient object detection has received great amount of attention in recent years. In this letter, we propose a novel salient object detection algorithm, which combines the global contextual information along with the low-level edge features. First, we train an edge detection stream based on the state-of-the-art holistically-nested edge detection (HED) model and extract hierarchical boundary information from each VGG block. Then, the edge contours are served as the complementary edge-aware information and integrated with the saliency detection stream to depict continuous boundary for salient objects. Finally, we combine pyramid pooling modules with auxiliary side output supervision to form the multi-scale pyramid-based supervision module, providing multi-scale global contextual information for the saliency detection network. Compared with the previous methods, the proposed network contains more explicit edge-aware features and exploit the multi-scale global information more effectively. Experiments demonstrate the effectiveness of the proposed method, which achieves the state-of-the-art performance on five popular benchmarks.
Wenlong Guan, Tiantian Wang 0002, Jinqing Qi, Lihe Zhang, Huchuan Lu
IEEE Signal Process. Lett.3
2018 Progressive Attention Guided Recurrent Network for Salient Object Detection
abstract
Effective convolutional features play an important role in saliency estimation but how to learn powerful features for saliency is still a challenging task. FCN-based methods directly apply multi-level convolutional features without distinction, which leads to sub-optimal results due to the distraction from redundant details. In this paper, we propose a novel attention guided network which selectively integrates multi-level contextual information in a progressive manner. Attentive features generated by our network can alleviate distraction of background thus achieve better performance. On the other hand, it is observed that most of existing algorithms conduct salient object detection by exploiting side-output features of the backbone feature extraction network. However, shallower layers of backbone network lack the ability to obtain global semantic information, which limits the effective feature learning. To address the problem, we introduce multi-path recurrent feedback to enhance our proposed progressive attention driven framework. Through multi-path recurrent connections, global semantic information from the top convolutional layer is transferred to shallower layers, which intrinsically refines the entire network. Experimental results on six benchmark datasets demonstrate that our algorithm performs favorably against the state-of-the-art approaches.
Tiantian Wang 0002, Jinqing Qi, Huchuan Lu, Gang Wang 0012
CVPR3
2018 Structured Siamese Network for Real-Time Visual Tracking
Lijun Wang 0001, Jinqing Qi, Dong Wang 0004, Mengyang Feng, Huchuan Lu
ECCV (9)3
2018 Predicting human gaze with multi-level information
Jinqing Qi, Huchuan Lu
Signal Process.4
2017 Saliency detection via joint modeling global shape and local consistency
Jinqing Qi, Shijing Dong, Huchuan Lu
Neurocomputing1
2017 Salient Object Detection via Multiple Instance Learning
abstract
Object proposals are a series of candidate segments containing objects of interest, which are taken as preprocessing and widely applied in various vision tasks. However, most of existing saliency approaches only utilize the proposals to compute a location prior. In this paper, we naturally take the proposals as the bags of instances of multiple instance learning (MIL), where the instances are the superpixels contained in the proposals, and formulate saliency detection problem as a MIL task (i.e., predict the labels of instances using the classifier in the MIL framework). This method allows some flexibility in finding a decision boundary based on the bag-level representations and can identify salient superpixels from ambiguous proposals. In addition, we introduce the MIL to an optimization mechanism, which iteratively updates training bags from easy to complex ones to learn a strong model. The significant improvement can be consistently achieved when applying the optimization model to existing saliency approaches. Extensive experiments demonstrate that the proposed algorithms perform favorably against the stateof- art saliency detection methods on several benchmark datasets.
Jinqing Qi, Huchuan Lu, Lihe Zhang, Xiang Ruan
IEEE Trans. Image Process.2
2017 Co-Bootstrapping Saliency
abstract
In this paper, we propose a visual saliency detection algorithm to explore the fusion of various saliency models in a manner of bootstrap learning. First, an original bootstrapping model, which combines both weak and strong saliency models, is constructed. In this model, image priors are exploited to generate an original weak saliency model, which provides training samples for a strong model. Then, a strong classifier is learned based on the samples extracted from the weak model. We use this classifier to classify all the salient and non-salient superpixels in an input image. To further improve the detection performance, multi-scale saliency maps of weak and strong model are integrated, respectively. The final result is the combination of the weak and strong saliency maps. The original model indicates that the overall performance of the proposed algorithm is largely affected by the quality of weak saliency model. Therefore, we propose a co-bootstrapping mechanism, which integrates the advantages of different saliency methods to construct the weak saliency model thus addresses the problem and achieves a better performance. Extensive experiments on benchmark data sets demonstrate that the proposed algorithm outperforms the state-of-the-art methods.
Huchuan Lu, Jinqing Qi, Na Tong, Xiang Ruan, Ming-Hsuan Yang 0001
IEEE Trans. Image Process.3
2016 Kernelized Subspace Ranking for Saliency Detection
Tiantian Wang 0002, Lihe Zhang, Huchuan Lu, Jinqing Qi
ECCV (8)5
2016 Saliency detection via a unified generative and discriminative model
Cong Jia, Jinqing Qi, Xiaohui Li 0005, Huchuan Lu
Neurocomputing2
2016 Salient object detection via point-to-set metric learning
Lihe Zhang, Jinqing Qi, Huchuan Lu
Pattern Recognit. Lett.3
2013 A Fast Approximate Sparse Coding Networks and Application to Image Denoising
Jianyong Cui, Jinqing Qi
ISNN (1)2