Jiqing Zhang

dblp:23/5881 · DBLP profile ↗
← Back
32ranked-venue papers
10as first author
27since 2021 · last 2026
0000-0002-0061-5465ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 8 first-author · 23 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 12 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Instance-Guided Scene Adaptation for Unsupervised Person Search
abstract
Unsupervised Domain Adaptation (UDA) is a challenging task in person search. It adapts a well-trained model from a labeled source domain to an unlabeled target domain for privacy and efficiency. Currently, most of the state-of-the-art UDA person search methods adopt multi-scale feature alignment techniques to learn domain-invariant representations. However, person search is a multi-granularity task, and such an indiscriminate method of bridging the differences between domains misleads the identity learning process, which significantly limits the model's performance. In this paper, we propose an Instance-Guided Scene Adaptation (IGSA) framework by eradicating scene disparities and focusing the tasks on instances, effectively eliminating the contradiction between person search and domain adaptation. In IGSA, a Scene-Aware Bidirectional Filter (SABF) is designed to divide the image features into background and foreground to perform bidirectional modulations, thereby achieving simultaneous scene elimination and instance enhancement. To further improve the reliability of identity learning, we also propose an Instance Consistency Contrastive Learning (ICCL) method. By performing cross-epoch updates on the instance-level memory bank and re-initializing the cluster-level memory bank, the problem of inconsistent training across epochs caused by instance identity drift can be alleviated. Through the above designs, our method can achieve state-of-the-art performance on two benchmark datasets, with 82.1% mAP and 83.8% top-1 on the CUHK-SYSU dataset and 41.1% mAP and 82.3% top-1 on the PRW dataset, which is even better than some supervised methods.
Huibing Wang, Jinjia Peng, Xianping Fu, Jiqing Zhang
AAAI5
2026 Localization-Anchored Instance Discrimination for Domain Adaptive Person Search
abstract
Domain-adaptive person search (DAPS) aims to transfer pedestrian detection and re-identification capabilities from a labeled source domain to an unlabeled target domain, yet faces critical challenges from domain shift: semantic confusion among overlapping instances, over-reliance on shallow features for look-alike targets, and poor discriminability of small-scale instances. To address these issues, we propose the Localization-Anchored Instance Discrimination (LAID) framework, which leverages spatial relationships between bounding boxes as auxiliary signals to enhance instance identity learning. LAID integrates three complementary strategies: 1) Cost-Aware Instance Matching (CAIM) uses IoU-based global optimal assignment to align current detections with historical identities, reducing overlap-induced misassociations; 2) Dual-Scope Contrastive Learning (DSCL) combines spatial separation constraints (for geometrically distant pairs) with global contrastive learning, prompting the model to learn deep discriminative features beyond superficial similarities; 3) Task-Sensitivity Alignment (TSA) aligns confidence distributions of detection and ReID heads via KL divergence, ensuring consistent pseudo-label generation. Extensive experiments on CUHK-SYSU and PRW datasets demonstrate that LAID outperforms state-of-the-art DAPS methods, validating its effectiveness in mitigating domain shift and narrowing the performance gap between supervised and domain-adaptive person search.
Linfeng Qi 0001, Huibing Wang, Jinjia Peng, Jiqing Zhang
AAAI4
2026 AlignTrack: Top-Down Spatiotemporal Resolution Alignment for RGB-Event Visual Tracking
Jiqing Zhang, Yang Wang 0106, Yuanchen Wang, Xin Yang 0011
AAAI2
2026 Efficient Vision Transformer with Token Sparsification for Event-Based Object Tracking
Jiqing Zhang, Xin Yang 0011, Haoming Tang, Yuanchen Wang, Huibing Wang, Xianping Fu
Int. J. Comput. Vis.1
2026 A Lightweight Polarization-Guided Plug-In for Underwater Image Enhancement
abstract
Underwater images play a vital role in marine exploration, but are often severely degraded due to complex imaging conditions, including color distortion, haze effects, and non-uniform illumination. Existing deep learning-based enhancement methods predominantly rely on conventional RGB sensors, which struggle to distinguish between scattered and reflected light, thereby limiting enhancement performance. Polarization imaging, with its capability to capture directional light information, offers promising potential for underwater image enhancement. In this paper, we propose a lightweight yet effective polarization feature extractor that captures global spatial cues from polarization images. Additionally, we design a polarization-guided feature integration module that adaptively enhances the representational capacity of RGB features. Notably, the proposed module is plug-in and can be seamlessly integrated into existing RGB-based enhancement networks. Extensive experiments across multiple datasets demonstrate that incorporating polarization information significantly improves enhancement performance, highlighting its effectiveness as a valuable cue for underwater image enhancement. The code and pretrained models are at https://github.com/jgy0/UPGD.
Guangyao Ju, Jiqing Zhang, Jingqi Zang, Zetian Mi, Xin Yang 0011, Huibing Wang, Jiarui Fan, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.2
2026 Lightweight and Personalized Single-Eye Emotion Recognition via CNN-SNN Spatiotemporal Learning and Memory-Inferred Event Features
abstract
Emotion recognition is essential for improving user experience and interaction quality in human-centered applications. While recent studies have leveraged both event and traditional cameras to enhance eye-based emotion recognition, their practical deployment is hindered by the scarcity of event cameras and the complexity of dual-modality frameworks. Personalization, which is critical for handling individual differences in emotional expression, is also affected by these factors, resulting in reduced performance and adaptation efficiency. To address these challenges, we propose a lightweight and personalized single-eye emotion recognition network, called LPSEER. LPSEER introduces a novel hybrid neural architecture that integrates a convolutional neural network (CNN) and a spiking neural network (SNN) to capture spatiotemporal features from video frames and events, respectively. Additionally, we design a memorybased event feature inference (MEFI) module that recalls event features from video frames, eliminating the reliance on event cameras during inference and personalization while retaining the discriminative advantages of event-based representations. Experimental results demonstrate that LPSEER achieves state-of- the-art recognition accuracy while maintaining the smallest model size and lowest computational cost. Further experiments confirm the strong generalization capabilities and the ability to achieve faster, more accurate personalization. These advantages collectively enable lightweight, accurate, and efficient emotion recognition for real-world human-centered applications.
Qianhui Liu, Jiqing Zhang, Yang Wang 0106, Malu Zhang, Xin Yang 0011, Gang Pan 0001, Haizhou Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 Interaction-Driven Edge Crisping for Underwater Salient Object Detection
abstract
Underwater salient object detection (USOD) faces greater challenges than general scenes due to the edge blurring which is caused by light absorption and scattering in water. Existing methods employ unrefined edge feature to perform unidirectional guidance on saliency feature, resulting in the coarse edge of saliency map. To address this issue, we propose a novel interaction-driven edge crisping network (IDENet) for underwater salient object detection. IDENet facilitates the bidi-rectional modulation of inter-features and the self-refinement of intra-feature, generates crisp saliency map and edge map. In IDENet, the interaction-driven edge guidance module (IDEGM) is designed to utilize cross-feature interaction by leveraging their correlations, facilitating saliency feature’s awareness of edge information, mitigating the interference of non-salient objects in edge feature. To learn more accurate edge region of the salient object, the edge intersection-and-union loss function (EIUL) is introduced to restrict the intersection and union of predicted saliency maps and edge maps to prevent over-expansion or under-contraction. Experimental results on two latest underwater datasets demonstrate the superiority of the proposed method over the state-of-the-art models. The source code of our method will be made available at https://github.com/ UnderwaterVisionMZTdlmu/IDENet.
Zetian Mi, Shuaiyong Jiang, Guanxi Li, Jiqing Zhang, Huibing Wang, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.5
2025 Unsupervised Domain Adaptive Person Search via Dual Self-Calibration
abstract
Unsupervised Domain Adaptive (UDA) person search focuses on employing the model trained on a labeled source domain dataset to a target domain dataset without any additional annotations. Most effective UDA person search methods typically utilize the ground truth of the source domain and pseudo-labels derived from clustering during the training process for domain adaptation. However, the performance of these approaches will be significantly restricted by the disrupting pseudo-labels resulting from inter-domain disparities. In this paper, we propose a Dual Self-Calibration (DSCA) framework for UDA person search that effectively eliminates the interference of noisy pseudo-labels by considering both the image-level and instance-level features perspectives. Specifically, we first present a simple yet effective Perception-Driven Adaptive Filter (PDAF) to adaptively predict a dynamic filter threshold based on input features. This threshold assists in eliminating noisy pseudo-boxes and other background interference, allowing our approach to focus on foreground targets and avoid indiscriminate domain adaptation. Besides, we further propose a Cluster Proxy Representation (CPR) module to enhance the update strategy of cluster representation, which mitigates the pollution of clusters from misidentified instances and effectively streamlines the training process for unlabeled target domains. With the above design, our method can achieve state-of-the-art (SOTA) performance on two benchmark datasets, with 80.2% mAP and 81.7% top-1 on the CUHK-SYSU dataset, with 39.9% mAP and 81.6% top-1 on the PRW dataset, which is comparable to or even exceeds the performance of some fully supervised methods.
Linfeng Qi 0001, Huibing Wang, Jiqing Zhang, Jinjia Peng
AAAI3
2025 Exploring Historical Information for RGBE Visual Tracking with Mamba
abstract
Combining the advantages of conventional and event cameras for robust visual tracing has drawn extensive interest. However, existing tracking approaches heavily engage in complex cross-modal fusion modules, leading to higher computational complexity and training challenges. Besides, these methods generally ignore the effective integration of historical information, which is crucial to grasping the change in the target’s appearance and motion trends. Given the recent advancements in Mamba’s long-range modeling and linear complexity, we explore its potential in addressing the above issues in RGBE tracking tasks. Specifically, we first propose an efficient fusion module based on Mamba, which utilizes a simple gate-based interaction scheme to achieve effective modality-selective fusion. This module can be seamlessly integrated into the encoding layer of prevalent Transformer-Based backbones. Moreover, we further present a novel historical decoder that leverages Mamba’s advanced long sequence modeling to effectively capture the target appearance changes with autoregressive queries. Extensive experiments show that our proposed approach achieves state-of-the-art performance on multiple challenging short-term and long-term RGBE benchmarks. Besides, the effectiveness of each key Mamba-Based component of our approach is evidenced by our thorough ablation study.Code will be released at: https://github.com/scy0712/MamTrack
Jiqing Zhang, Huilin Ge, Qianchen Xia
CVPR2
2025 Scalable Multi-view Clustering based on Tight Anchor Distribution
Yawei Chen, Huibing Wang, Mingze Yao, Jinjia Peng, Guangqi Jiang, Jiqing Zhang
ACM Multimedia6
2025 Dual-Constraint Multi-view Fuzzy Clustering with Scalable Anchor Graph Learning
Luyan Cui, Huibing Wang, Yawei Chen, Mingze Yao, Xianping Fu, Jiqing Zhang
ACM Multimedia6
2025 Eye-based Emotion Recognition via Event-Driven Sparse Transformers
abstract
Event-driven eye-based emotion recognition has attracted increasing attention due to the high temporal resolution and dynamic range inherent to event cameras. The intrinsic spatial sparsity of event data, combined with the eye-based emotion recognition task's reliance on localized features such as eyebrows and eyelids, makes it intuitive and efficient to discard less informative regions. However, integrating such sparsification into CNNs remains challenging due to their reliance on dense grid-based operations. In this paper, we propose an efficient vision transformer framework for eye-based emotion recognition with event cameras. Specifically, we present window selection and token selection schemes tailored for event data and eye-based emotion recognition, which can diminish computing demands while enhancing performance. Firstly, we estimate the importance of all local windows and discard those with limited information, reducing computational cost while emphasizing attention on the periocular region. Secondly, we further introduce an adaptive token pruning mechanism that jointly evaluates the input event data and tokens to predict a binary decision mask, identifying and discarding uninformative tokens. Extensive experiments validate that the proposed approach outperforms existing state-of-the-art methods in accuracy by a significant margin.
Zixuan Wan, Jiqing Zhang, Yafei Wang 0004, Zetian Mi, Xin Yang 0011, Xianping Fu, Huibing Wang
ACM Multimedia2
2025 Fusion-Based Channel-Wise Isotropic Convergent Real-Time Underwater Image Enhancement
abstract
Existing underwater image enhancement (UIE) methods typically prioritize improving image quality at the expense of algorithmic efficiency. In this paper, we propose a fusion-based, channel-wise isotropic convergent UIE method designed for real-time performance. The proposed approach comprises three key modules: (i) a non-linear transformation module that corrects color casts and aligns the pixel distribution with the gray-world assumption (GWA); (ii) a channel-wise isotropic convergence scheme that reduces intensity distribution disparities across channels, promoting balanced convergence; and (iii) a patch-based enhancement strategy that divides the image into smaller patches to better capture local features and improve adaptability to non-uniform degradation. Moreover, certain critical steps in our method are optimized to achieve O(1) time complexity, allowing it to meet real-time requirements. Extensive experiments validate the effectiveness of each module in the proposed method, showcasing its superiority when compared to the existing state-of-the-art (SOTA) approaches. Code has been released at https://github.com/JohnChenS/FCICE_UnderwaterImageEnhancement.
Yuehan Chen, Jiqing Zhang, Haoming Tang, Huibing Wang, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.2
2025 An Underwater Image Restoration Method With Polarization Imaging Optimization Model for Poor Visible Conditions
abstract
Polarization imaging is extensively employed in underwater image restoration due to its effectiveness in removing backscattered light. However, existing polarization imaging methods generally assume the degree of polarization (DoP) of the backscattering is spatially constant and estimate it from the background region, limiting their practical applications. To address these challenges, we propose an underwater image restoration method based on a polarization imaging optimization model (PIOM). First, we develop a novel polarization image formation model by fusing the DoP and angle of polarization (AoP) of backscattered light. Second, we introduce an adaptive particle swarm local optimization (APSLO) method based on the PIOM. This method decomposes the image into small blocks and employs an objective optimization function to estimate the local optimal fusion parameters. Additionally, we propose a robust polynomial spatial fitting method to reduce block artifacts and noise disturbances, achieving globally optimal fusion parameters. Finally, we fully consider the advantages of gamma correction, and propose an adaptive contrast enhancement method to balance brightness and contrast. Experimental results show that our PIOM effectively removes backscattering while preserving finer details, colors, and contours. The code and datasets will be available athttps://github.com/liyafengLYF/UIRPIOM.
Yuehan Chen, Jiqing Zhang, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.3
2025 Underwater Vignetting Image Correction Based on Binary Polynomial Regularization and Latent Low-Rank Representation
abstract
Due to light attenuation and complex environments, underwater robots need to carry artificial light to improve visibility, which leads to issues such as brightness vignetting, low contrast, and color distortion in the captured underwater images. However, existing methods for enhancing underwater images often overlook the challenges caused by artificial light. To address these challenges, we construct a novel underwater vignetting image formation model and propose a correction method called UVIC. The method consists of three main modules: separating the vignetting component, separating the backscattering component, and adaptive brightness and color correction. In our model, based on the linear relationship between the image gradient and the coefficients of the binary polynomial, we introduce a binary polynomial regularization to separate the vignetting component without estimating the center of the vignetting. Additionally, the backscattering can be effectively separated by introducing a latent low-rank representation based on local consistency, without estimating atmospheric light and transmission parameters. Furthermore, we design an adaptive brightness and color correction module using the global brightness of the image L layer and the histogram distribution characteristics of the a and b layers to adjust the brightness and color bias of the image. Particularly, there are both additive and multiplicative operations, and we decompose the objective function into two submodels and solve them by the iterative reweighted least squares and alternating direction multiplier methods, respectively. Numerous experiments demonstrate that UVIC not only effectively corrects image brightness vignetting, but also improves color bias, contrast, and sharpness.
Yulin Wang 0003, Yueming Ma, Jiqing Zhang, Zetian Mi, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.4
2025 DG-SMOTE: A Distance-Angle-Based Genetic Synthetic Minority Over-Sampling Technique for Unbalanced Data Learning
abstract
Many real-world applications often generate unbalanced data. Learning from such data may lead to biased classifiers that perform poorly on the class of interest. Oversampling methods have been shown to be effective in rebalancing unbalanced data to help classifiers avoid performance bias. However, many existing oversampling methods rely on a predesigned linear model structure and the neighborhood information of an original instance. This may lead to the generation of noisy instances when the original data has noise. In this study, we develop a novel oversampling method in which genetic programming is introduced to automatically select good-quality instances and evolve a model structure that combines the selected instances to create a new instance. In the proposed oversampling method, an individual is used to represent a generated instance, which is evaluated by the fitness function designed based on the Euclidean distance and the cosine theorem. In the experiments, we examine the effectiveness of the proposed oversampling method in assisting different types of classifiers to solve the issue of class imbalance, and compare it with popular sampling methods in unbalanced classification. The results have been analyzed comprehensively, indicating that the new method successfully addressed the class imbalance issue by generating a group of good-quality instances for the minority class and outperformed the compared sampling methods in almost all cases.
Wenbin Pei, Yuyang Cui, Bing Xue 0001, Mengjie Zhang 0001, Jiqing Zhang, Yaqing Hou, Guangyu Zou, Qiang Zhang 0008
IEEE Trans. Evol. Comput.5
2025 Spiking Neural Networks With Adaptive Membrane Time Constant for Event-Based Tracking
abstract
The brain-inspired Spiking Neural Networks (SNNs) work in an event-driven manner and have an implicit recurrence in neuronal membrane potential to memorize information over time, which are inherently suitable to handle temporal event-based streams. Despite their temporal nature and recent approaches advancements, these methods have predominantly been assessed on event-based classification tasks. In this paper, we explore the utility of SNNs for event-based tracking tasks. Specifically, we propose a brain-inspired adaptive Leaky Integrate-and-Fire neuron (BA-LIF) that can adaptively adjust the membrane time constant according to the inputs, thereby accelerating the leakage of meaningless noise features and reducing the decay of valuable information. SNNs composed of our proposed BA-LIF neurons can achieve high performance without a careful and time-consuming trial-by-error initialization on the membrane time constant. The adaptive capability of our network is further improved by introducing an extra temporal feature aggregator (TFA) that assigns attention weights over the temporal dimension. Extensive experiments on various event-based tracking datasets validate the effectiveness of our proposed method. We further validate the generalization capability of our method by applying it to other event-classification tasks.
Jiqing Zhang, Malu Zhang, Yuanchen Wang, Qianhui Liu, Haizhou Li 0001, Xin Yang 0011
IEEE Trans. Image Process.1
2025 Dark Channel Low-Rank Prior for Enhanced Single Underwater Image Restoration
Yulin Wang 0003, Zheng Liang 0001, Zetian Mi, Jiqing Zhang, Xianping Fu
Vis. Comput.4
2024 Adaptive Vision Transformer for Event-Based Human Pose Estimation
abstract
Event-based human pose estimation has gained popularity due to the benefits of high temporal resolution and high dynamic range offered by event cameras. The inherent spatial sparsity of event data makes discarding less significant regions a straightforward and effective way to decrease the computation. However, implementing this operation in CNNs poses a challenge, as it disrupts the regularity of dense convolutional workload. In this paper, we propose an adaptive vision transformer, a novel efficient backbone for human pose estimation with event cameras. Specifically, we present two adaptive patch and token sampling approaches based on the characteristics of events, thereby reducing the computational load while still achieving comparable performance. Firstly, we design an adaptive patch sampling scheme to eliminate inactivity patches by assessing the entropy of the events before they are inputted into the transformer. Secondly, we further propose an adaptive token reduction strategy to selectively remove less informative tokens in transformer layers through a dynamic token pruning algorithm. To exploit event-based visual cues in human pose estimation tasks, we construct a large-scale frame-event-based dataset, dubbed Event Multi Movement HPE (EventMM HPE). The dataset provides annotation frequencies up to 240 Hz. Extensive experiments demonstrate that our proposed approach outperforms existing state-of-the-art methods in estimation accuracy. The source code and dataset are available at https://github.com/doublemanyu/Adaptive-Vision-Transformer-for-Event-Based-HPE.
Nannan Yu, Jiqing Zhang, Yuji Zhang 0004, Qirui Bao, Xiaopeng Wei, Xin Yang 0011
ACM Multimedia3
2024 A Universal Event-Based Plug-In Module for Visual Object Tracking in Degraded Conditions
Jiqing Zhang, Bo Dong 0004, Yingkai Fu, Yuanchen Wang, Xiaopeng Wei, Xin Yang 0011
Int. J. Comput. Vis.1
2024 Underwater image restoration via spatially adaptive polarization imaging and color correction
Jiqing Zhang, Yuehan Chen, Haoming Tang, Xianping Fu
Knowl. Based Syst.2
2023 Frame-Event Alignment and Fusion Network for High Frame Rate Tracking
abstract
Most existing RGB-based trackers target low frame rate benchmarks of around 30 frames per second. This setting restricts the tracker's functionality in the real world, especially for fast motion. Event-based cameras as bioinspired sensors provide considerable potential for high frame rate tracking due to their high temporal resolution. However, event-based cameras cannot offer fine-grained texture information like conventional cameras. This unique complementarity motivates us to combine conventional frames and events for high frame rate object tracking under various challenging conditions. In this paper, we propose an end-to-end network consisting of multi-modality alignment and fusion modules to effectively combine meaningful information from both modalities at different measurement rates. The alignment module is responsible for cross-style and cross-frame-rate alignment between frame and event modalities under the guidance of the moving cues furnished by events. While the fusion module is accountable for emphasizing valuable features and suppressing noise information by the mutual complement between the two modalities. Extensive experiments show that the proposed approach outper-forms state-of-the-art trackers by a significant margin in high frame rate tracking. With the FE240Hz dataset, our approach achieves high frame rate tracking up to 240Hz.
Jiqing Zhang, Yuanchen Wang, Wenxi Liu, Meng Li 0072, Jinpeng Bai, Xin Yang 0011
CVPR1
2023 Distractor-Aware Event-Based Tracking
abstract
Event cameras, or dynamic vision sensors, have recently achieved success from fundamental vision tasks to high-level vision researches. Due to its ability to asynchronously capture light intensity changes, event camera has an inherent advantage to capture moving objects in challenging scenarios including objects under low light, high dynamic range, or fast moving objects. Thus event camera are natural for visual object tracking. However, the current event-based trackers derived from RGB trackers simply modify the input images to event frames and still follow conventional tracking pipeline that mainly focus on object texture for target distinction. As a result, the trackers may not be robust dealing with challenging scenarios such as moving cameras and cluttered foreground. In this paper, we propose a distractor-aware event-based tracker that introduces transformer modules into Siamese network architecture (named DANet). Specifically, our model is mainly composed of a motion-aware network and a target-aware network, which simultaneously exploits both motion cues and object contours from event data, so as to discover motion objects and identify the target object by removing dynamic distractors. Our DANet can be trained in an end-to-end manner without any post-processing and can run at over 80 FPS on a single V100. We conduct comprehensive experiments on two large event tracking datasets to validate the proposed model. We demonstrate that our tracker has superior performance against the state-of-the-art trackers in terms of both accuracy and efficiency.
Yingkai Fu, Meng Li 0072, Wenxi Liu, Yuanchen Wang, Jiqing Zhang, Xiaopeng Wei, Xin Yang 0011
IEEE Trans. Image Process.5
2022 Spiking Transformers for Event-based Single Object Tracking
abstract
Event-based cameras bring a unique capability to tracking, being able to function in challenging real-world conditions as a direct result of their high temporal resolution and high dynamic range. These imagers capture events asynchronously that encode rich temporal and spatial information. However, effectively extracting this information from events remains an open challenge. In this work, we propose a spiking transformer network, STNet, for single object tracking. STNet dynamically extracts and fuses information from both temporal and spatial domains. In particular, the proposed architecture features a transformer module to provide global spatial information and a spiking neural network (SNN) module for extracting temporal cues. The spiking threshold of the SNN module is dynamically adjusted based on the statistical cues of the spatial information, which we find essential in providing robust SNN features. We fuse both feature branches dynamically with a novel cross-domain attention fusion algorithm. Extensive experiments on three event-based datasets, FE240hz, EED and VisEvent validate that the proposed STNet outperforms existing state-of-the-art methods in both tracking accuracy and speed with a significant margin. The code and pretrained models are at https://github.com/Jee-King/CVPR2022_STNet.
Jiqing Zhang, Bo Dong 0004, Haiwei Zhang 0003, Jianchuan Ding, Felix Heide, Xin Yang 0011
CVPR1
2022 A Two-Stage Attentive Network for Single Image Super-Resolution
abstract
Recently, deep convolutional neural networks (CNNs) have been widely explored in single image super-resolution (SISR) and contribute remarkable progress. However, most of the existing CNNs-based SISR methods do not adequately explore contextual information in the feature extraction stage and pay little attention to the final high-resolution (HR) image reconstruction step, hence hindering the desired SR performance. To address the above two issues, in this paper, we propose a two-stage attentive network (TSAN) for accurate SISR in a coarse-to-fine manner. Specifically, we design a novel multi-context attentive block (MCAB) to make the network focus on more informative contextual features. Moreover, we present an essential refined attention block (RAB) which could explore useful cues in HR space for reconstructing fine-detailed HR image. Extensive evaluations on four benchmark datasets demonstrate the efficacy of our proposed TSAN in terms of quantitative metrics and visual effects. Code is available athttps://github.com/Jee-King/TSAN.
Jiqing Zhang, Chengjiang Long, Yuxin Wang 0001, Haiyin Piao, Haiyang Mei, Xin Yang 0011
IEEE Trans. Circuits Syst. Video Technol.1
2021 Object Tracking by Jointly Exploiting Frame and Event Domain
abstract
Inspired by the complementarity between conventional frame-based and bio-inspired event-based cameras, we propose a multi-modal based approach to fuse visual cues from the frame- and event-domain to enhance the single object tracking performance, especially in degraded conditions (e.g., scenes with high dynamic range, low light, and fast-motion objects). The proposed approach can effectively and adaptively combine meaningful information from both domains. Our approach’s effectiveness is enforced by a novel designed cross-domain attention schemes, which can effectively enhance features based on self- and cross-domain attention schemes; The adaptiveness is guarded by a specially designed weighting scheme, which can adaptively balance the contribution of the two domains. To exploit event-based visual cues in single-object tracking, we construct a large-scale frame-event-based dataset, which we subsequently employ to train a novel frame-event fusion based model. Extensive experiments show that the proposed approach outperforms state-of-the-art frame-based tracking methods by at least 10.4% and 11.9% in terms of representative success rate and precision rate, respectively. Besides, the effectiveness of each key component of our approach is evidenced by our thorough ablation study.
Jiqing Zhang, Xin Yang 0011, Yingkai Fu, Xiaopeng Wei, Bo Dong 0004
ICCV1
2021 Multi-domain collaborative feature representation for robust visual object tracking
Jiqing Zhang, Bo Dong 0004, Yingkai Fu, Yuxin Wang 0001, Xin Yang 0011
Vis. Comput.1
2020 Multi-Context And Enhanced Reconstruction Network For Single Image Super Resolution
abstract
Most existing single image super-resolution (SISR) methods continually increase the depth or width of networks, without adequately exploring contextual features which are essential for reconstruction. Moreover, such existing methods pay little attention to the final high-resolution(HR) image reconstruction step and therefore hinder the desired SR performance. In this paper, we propose a multi-context and enhanced reconstruction network (MCERN) for SISR. Specifically, a novel model named Multi-Context Block (MCB) which extracts more image contextual features with multibranch dilated convolution. Applying multiple MCBs with residual and dense connections, we can effectively extract contextual and hierarchical features for obtaining the coarse super-resolution result. Then an enhanced reconstruction block (ERB) is followed to extract essential spatial features on the high-resolution image to refine the coarse result to a better result. Extensive benchmark evaluations demonstrate the efficacy of our proposed MCERN in terms of metric accuracy and visual effects.
Jiqing Zhang, Chengjiang Long, Yuxin Wang 0001, Xin Yang 0011, Haiyang Mei
ICME1
2019 Two-Dimensional Calibration for Fixed-Pattern Noise Reduction of Thermal Images
abstract
A two-dimensional calibration technique is proposed to reduce fixed-pattern noise in infrared thermal images for extending dynamic range of sampling integration time. Traditional two-point calibration generates correction coefficients by irradiance of two reference temperatures, providing excellent spatial noise performance for the predetermined sampling integration time. To break the limitation of invariable integration time, two-dimensional calibration is proposed, which generates correction coefficients not only under irradiance of two temperatures, but also with two integration times. By exploring information from both temperature and integration, the new approach can reduce fixed-pattern noise for different sampling integration time, requiring same computation complexity and hardware consumption as conventional two-point calibration. In experiments with infrared camera, the proposed method suppresses fixed-pattern noise to a low level with same correction coefficients for sampling integration time changing from 0.4 ms to 1.6 ms. Compared to traditional calibration-based approaches, the presented scheme extends dynamic range of integration time by four times.
Nan Chen 0003, Jiqing Zhang, Shengyou Zhong, Wenbiao Mao, Douming Hu, Libin Yao
ISCAS2
2019 DRFN: Deep Recurrent Fusion Network for Single-Image Super-Resolution With Large Factors
abstract
Recently, single-image super-resolution has made great progress due to the development of deep convolutional neural networks (CNNs). The vast majority of CNN-based models use a predefined upsampling operator, such as bicubic interpolation, to upscale input low-resolution images to the desired size and learn nonlinear mapping between the interpolated image and ground truth high-resolution (HR) image. However, interpolation processing can lead to visual artifacts as details are over smoothed, particularly when the super-resolution factor is high. In this paper, we propose a deep recurrent fusion network (DRFN), which utilizes transposed convolution instead of bicubic interpolation for upsampling and integrates different-level features extracted from recurrent residual blocks to reconstruct the final HR images. We adopt a deep recurrence learning strategy and, thus, have a larger receptive field, which is conducive to reconstructing an image more accurately. Furthermore, we show that the multilevel fusion structure is suitable for dealing with image super-resolution problems. Extensive benchmark evaluations demonstrate that the proposed DRFN performs better than most current deep learning methods in terms of accuracy and visual effects, especially for large-scale images, while using fewer parameters.
Xin Yang 0011, Haiyang Mei, Jiqing Zhang, Ke Xu 0010, Qiang Zhang 0008, Xiaopeng Wei
IEEE Trans. Multim.3
2012 Gesture recognition using video and floor pressure data
abstract
This paper presents a multimodal gesture recognition framework using video and floor pressure data. The key contribution of this research is to show that using additional floor pressure data significantly improves the recognition of visually ambiguous gestures. To effectively combine gesture recognition results from both the visual and pressure sensing modalities, we have adopted a two-stage cascaded sequential information integration scheme. In Stage-1 of the scheme, an unknown movement segment is first classified into a gesture group based on the visual features, and then in Stage-2, the input movement is further recognized as a gesture within the gesture group according to the pressure features. In the proposed framework, the hidden Markov models (HMMs) are used to model and recognize gestures using features from video and pressure data. The experimental results obtained on an in-house video and floor pressure gesture dataset demonstrate the efficacy of the proposed multimodal gesture recognition framework.
Gang Qian, Bo Peng 0005, Jiqing Zhang
ICIP3
2009 Footprint tracking and recognition using a pressure sensing floor
abstract
This paper presents an approach to clustering, tracking and recognizing footprints of a single subject from pressure data obtained using a pressure sensing floor. The proposed method clusters active footprint areas on the floor and recognizes and tracks the footprints according to their 2D shapes and geometrical relationships among foot clusters. Experimental results show the efficacy of the proposed approach.
Jiqing Zhang, Gang Qian, Assegid Kidané
ICIP1