Qijun Zhao

dblp:26/569 · DBLP profile ↗
← Back
155ranked-venue papers
13as first author
104since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 92 · 6 first-author · 63 since 2021Artificial intelligence and machine learning · 87 · 11 first-author · 52 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 5 since 2021Security and privacy · 7 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DiAPR: Dimensionally-Allocated Prototype Refinement for Non-Exemplar Class Incremental Learning
abstract
Non-Exemplar Class Incremental Learning (NECIL) strives to preserve classification performance in an evolving data stream without revisiting old-class exemplars. Current methods mitigate catastrophic forgetting by replaying and augmenting historical prototypes as surrogates for old classes. However, they treat prototypes as holistic representations for global-level augmentations, which overlook dimensional semantic disparity and old-new class relationships, failing to maintain old-class discriminability and adaptability to the evolving feature space. To address this challenge, we propose Dimensionally-Allocated Prototype Refinement (DiAPR), a granular framework that progressively refines prototypes to exhibit class separability in the new feature space through three modules. Specifically, Distribution-aware Pairing (DAP) captures old-new class semantic consistency to guide Granular Semantic Allocation (GSA) in dimension-wise conflation, while Cross-Dimensional Transition (CDT) enhances cross-dimensional dependencies. The resulting prototypes sharpen classifier decision boundaries. Moreover, CDT inherently enables softened feature alignment, thereby yielding a more compatible feature space. Extensive experiments demonstrate DiAPR’s superiority, with improvements over SOTA by 2.35%, 0.70%, 0.96% on three CIFAR-100 settings, 1.03%, 0.54%, 0.40% on Tiny-ImageNet, and 0.60% on ImageNet-Subset.
Ruixuan Gao, Qijun Zhao, Keren Fu
AAAI2
2026 SAR-DisentDM: A Semantic-Disentangled Diffusion Model for Limited-Data SAR Image Synthesis
abstract
The high cost of synthetic aperture radar (SAR) data acquisition motivates SAR image generation research. However, the data scarcity and SAR's inherent azimuth sensitivity make generative models suffer from severe azimuth overfitting. Most existing methods require supplementary data to work effectively, limiting their practicality. In this paper, we propose SAR-DisentDM, a novel semantic-disentangled diffusion model for limited-data SAR image generation, without requiring any auxiliary resources. We develop a physics-aware diffusion architecture that explicitly models semantic knowledge of SAR images, including intrinsic characteristics, contextual diversity, and measurement randomness. A key innovation is the attention-guided semantic disentanglement (AGSD) module, designed to decouple category-specific features from azimuth-variable scattering patterns. This is achieved by aid of a dual disentangled loss with time-step-adaptive optimization. Furthermore, we introduce an azimuth angle perturbation augmentation (AAPA) mechanism, to enhance the model's robustness to minor azimuth angle errors. Extensive evaluations validate that SAR-DisentDM enables controllable SAR image synthesis with designated attributes, significantly improving representation and generalization abilities under limited data. Synthetic imagery from our approach boosts automatic target recognition (ATR) accuracy beyond state-of-the-art methods.
Qijun Zhao, Zijian Deng
AAAI3
2026 CapeNext: Rethinking and Refining Dynamic Support Information for Category-Agnostic Pose Estimation
abstract
Recent research in Category-Agnostic Pose Estimation (CAPE) has adopted fixed textual keypoint description as semantic prior for two-stage pose matching frameworks. While this paradigm enhances robustness and flexibility by disentangling the dependency of support images, our critical analysis reveals two inherent limitations of static joint embedding: (1) polysemy-induced cross-category ambiguity during the matching process(e.g., the concept "leg" exhibiting divergent visual manifestations across humans and furniture), and (2) insufficient discriminability for fine-grained intra-category variations (e.g., posture and fur discrepancies between a sleeping white cat and a standing black cat). To overcome these challenges, we propose a new framework that innovatively integrates hierarchical cross-modal interaction with dual-stream feature refinement, enhancing the joint embedding with both class-level and instance-specific cues from textual description and specific images. Experiments on the MP-100 dataset demonstrate that, regardless of the network backbone, CapeNext consistently outperforms state-of-the-art CAPE methods by a large margin.
Dan Zeng 0002, Shuiwang Li, Qijun Zhao, Qiaomu Shen, Bo Tang 0016
AAAI4
2026 PKTA: Part-oriented Knowledge Transfer and Acquisition for Non-Exemplar Lifelong Person Re-Identification
Ruixuan Gao, Qijun Zhao, Yangqianqian Chen
FG2
2026 Two-Stage MAE with Dual-Asymmetry Learning for 3D Facial Paralysis Grading
Chengchao Li, Qijun Zhao, Shune Tan, Lanlan Wang, Chunlin Zhu, Chenman Zhang, Bingyu Chen 0008, Yan Ai
FG4
2026 Generic Deepfake Feature Space Discovery via Coarse-to-Fine Disentanglement Learning
Zongyong Deng, Sirui Zhou, Qijun Zhao, Lanfei Qiao
FG4
2026 Learning motion blur robust vision transformers for real-time UAV tracking
You Wu 0009, Xucheng Wang, Dan Zeng 0002, Hengzhou Ye, Xiaolan Xie 0002, Qijun Zhao, Shuiwang Li
Expert Syst. Appl.6
2026 Handwritten text line segmentation with TextSAM: An enhanced segment anything model via multi-module fusion
Yunjie Xiang, Yukai Xian, Xianmu Cairang, Dorji Gesang, Pubu Danzeng, Qijun Zhao
Int. J. Document Anal. Recognit.6
2026 Mutual adversarial attack-based adaptive blending for face morphing
Zongyong Deng, Qiaoyun He, Qijun Zhao, Zuyuan He, Lanfei Qiao
Neurocomputing3
2026 UniTrack: Unifying day and night tracking with continual learning
Jiwei Mo, Feixiang He, Pengzhi Zhong, Qijun Zhao, Dan Zeng 0002, Shuiwang Li, Xianhao Shen
Inf. Sci.4
2026 CC-GFRT: Class Correlation-based Granular Feature Refinement and Transfer for Non-Exemplar Class Incremental Learning
Ruixuan Gao, Zongyong Deng, Yue Yang 0026, Qijun Zhao
Knowl. Based Syst.4
2026 Unleashing the power of motion and depth: A selective fusion strategy for RGB-D video salient object detection
Daerji Suolang, Keren Fu, Qijun Zhao
Knowl. Based Syst.4
2026 Star-vmamba: VMamba-based spatiotemporal adaptive refinement network for remote sensing image change detection
Ying Pu, Qijun Zhao
Mach. Vis. Appl.3
2026 Hyperanimal: Identity hypersphere guided synthetic datasets generation for individual animal identification
Yue Yang 0026, Zongming Peng, Zongyong Deng, Qijun Zhao
Pattern Recognit.6
2026 Learning an Adaptive and View-Invariant Vision Transformer for Real-Time UAV Tracking
abstract
Transformer-based models have improved visual tracking, but most still cannot run in real time on resource-limited devices, especially for unmanned aerial vehicle (UAV) tracking. To achieve a better balance between performance and efficiency, we propose AVTrack, an adaptive computation tracking framework that adaptively activates transformer blocks through an Activation Module (AM), which dynamically optimizes the ViT architecture by selectively engaging relevant components. To address extreme viewpoint variations, we propose to learn view-invariant representations via mutual information (MI) maximization. In addition, we propose AVTrack-MD, an enhanced tracker incorporating a novel MI maximization-based multi-teacher knowledge distillation framework. Leveraging multiple off-the-shelf AVTrack models as teachers, we maximize the MI between their aggregated softened features and the corresponding softened feature of the student model, improving the generalization and performance of the student, especially under noisy conditions. Extensive experiments show that AVTrack-MD achieves performance comparable to AVTrack’s performance while reducing model complexity and boosting average tracking speed by over 17%. Codes is available at https://github.com/wuyou3474/AVTrack.
You Wu 0009, Xucheng Wang, Xiangyang Yang 0001, Hengzhou Ye, Dan Zeng 0002, Qijun Zhao, Shuiwang Li
IEEE Trans. Circuits Syst. Video Technol.8
2026 Toward Real-Time UAV Tracking With Adaptive and Background-Aware Vision Transformers
abstract
In Unmanned Aerial Vehicle (UAV) tracking, discriminative correlation filters (DCF) are popular for their speed, making them ideal for real-time use with limited resources. Recently, lightweight convolutional neural networks (CNNs) have offered a new approach. Through filter pruning, these CNNs maintain high accuracy and efficiency, making them a strong alternative, especially for greater precision. Despite these advancements, the potential of pure vision transformers (ViTs) in UAV tracking remains largely untapped, especially based on the paradigm of conditional computation. In this work, we introduce an adaptive and background-aware Vision Transformer (Aba-ViT) and leverage it to develop a real-time UAV tracking framework called Aba-ViTrack. The proposed Aba-ViT exploits an adaptive and background-aware token computation method to reduce inference time. This approach adaptively discards tokens based on learned halting probabilities, which a priori are higher for background tokens than target ones. To further improve efficiency, this paper proposes a novel classroom-style learning (CSL) approach, where robustness knowledge is transmitted vertically from teacher to students, and generalization capability is enhanced through horizontal mutual learning among students. This method is used to compress Aba-ViTrack, resulting in Aba-ViTrack++. The upgraded version achieves a better balance between accuracy and efficiency in real-time UAV tracking. This version achieves a better balance between accuracy and efficiency for real-time UAV tracking. Extensive experiments on six UAV tracking benchmarks demonstrate that the proposed method achieves state-of-the-art performance in UAV tracking. The code is available at https://github.com/xyyang317/Aba-ViTrack.
Xiangyang Yang 0001, Dan Zeng 0002, Xucheng Wang, Hengzhou Ye, Xiaolan Xie 0002, Qijun Zhao, Jihua Zhu, Shuiwang Li
IEEE Trans. Circuits Syst. Video Technol.6
2025 CAM-YOLO: A Multi-Scale Cultural Element Detection Network for Thangka Images
abstract
To fully understand the cultural content of Thangka images and solve the issues of a large number of parameters, low accuracy and slow inference speed of multi-scale cultural element detection model for Thangka images in complex background. In this paper, we propose a multi-scale cultural element detection model for Thangka images based on improved yolov8. Firstly, C2f-StarNB is integrated into the backbone network instead of the C2f module, and element multiplication is used to achieve the reduction of parameter computation and improvement of inference speed without the need to widen the network. Secondly, the convolutional attention and multi-scale feedforward network module (CAMS) is integrated after the neck network to effectively capture the global contextual information and local features of the Thangka image, and improve the multi-scale cultural element detection performance of the Thangka, especially the small cultural element detection performance is improved significantly. Finally, a more focused bounding box loss Focaler-DIoU is proposed to improve the difficult sample detection performance under sample imbalance conditions. Validated on the Thangka detection dataset, the model outperforms other state-of-the-art methods in terms of the number of parameters and detection accuracy, and achieves an accuracy of 95.6 on mAP@50. The experimental results show that our proposed method has competitive performance for multi-scale cultural element detection in Thangka images.
Jianjun Xia, Dingguo Gao, Qijun Zhao, Songtao Xu, Te Shen
CSCWD3
2025 Samba: A Unified Mamba-based Framework for General Salient Object Detection
abstract
Existing salient object detection (SOD) models primarily resort to convolutional neural networks (CNNs) and Transformers. However, the limited receptive fields of CNNs and quadratic computational complexity of transformers both constrain the performance of current models on discovering attention-grabbing objects. The emerging state space model, namely Mamba, has demonstrated its potential to balance global receptive fields and computational complexity. Therefore, we propose a novel unified framework based on the pure Mamba architecture, dubbed saliency Mamba (Samba), to flexibly handle general SOD tasks, including RGB/RGB-D/RGB-T SOD, video SOD (VSOD), and RGB-D VSOD. Specifically, we rethink Mamba’s scanning strategy from the perspective of SOD, and identify the importance of maintaining spatial continuity of salient patches within scanning sequences. Based on this, we propose a saliency-guided Mamba block (SGMB), incorporating a spatial neighboring scanning (SNS) algorithm to preserve spatial continuity of salient patches. Additionally, we propose a context-aware upsampling (CAU) method to promote hierarchical feature alignment and aggregations by modeling contextual dependencies Experimental results show that our Samba outperforms existing methods across five SOD tasks on 21 datasets with lower computational cost, confirming the superiority of introducing Mamba to the SOD areas. Our code is available at https://github.com/Jia-Hao999/Samba.
Keren Fu, Xiaohong Liu 0001, Qijun Zhao
CVPR4
2025 Camouflaged Object Detection via Neural Architecture Search
abstract
The core challenge in camouflaged object detection (COD) is identifying objects that blend seamlessly with their surroundings. Existing methods emulate the strategies biological organisms break camouflage by manually constructing modules with expert knowledge from existing segmentation tasks, making it difficult to accurately understand complex and unique camouflage semantics. We are the first to apply neural architecture search (NAS) to COD, introducing an automatic localization and refinement network called ALRNet. It explores a large search space to discover more effective camouflage-specific modules. Specifically, we propose a search-based automatic receptive field block (ARFB) to adaptively excavate hierarchical discriminative cues and decouple features in a multi-branch architecture. Moreover, we introduce an edge-assisted explicit and implicit refinement (EEIR) module, combining explicit priors with implicit search to create a dual-task structure for edge and segmentation knowledge interaction. Experiments on four benchmarks demonstrate that ALRNet outperforms 15 state-of-the-art methods. Codes are available at https://github.com/BoydeLi/ALRNet.
Keren Fu, Qijun Zhao
ICASSP3
2025 Lightweight Multi-Frequency Enhancement Network for RGB-D Video Salient Object Detection
abstract
RGB-D Video Salient Object Detection has gained increasing interest, but existing models often struggle to balance efficiency and accuracy, hindering their applications on resource-constrained devices. A key challenge in designing lightweight models is maintaining accuracy while reducing parameters. To address this issue and bridge the gap in lightweight RGB-D VSOD research, we propose a lightweight network architecture using MobileNetV2 as the backbone. We introduce an Improved Cross-Shift Module (ICSM) to extract the fused depth and flow features with minimal overhead and a Multi-Frequency Enhancement Module (MFEM) to separate high-and low-frequency information and enhance the resulting feature maps using different techniques for final saliency prediction. Experimental results demonstrate that our method achieves competitive accuracy compared to non-efficient models, running at 80 FPS on a GPU with only 4.75M parameters, making it suitable for real-time applications. Code will be available at https://github.com/Tibetsonam/MFENet.
Daerji Suolang, Wangchuk Tsering, Keren Fu, Qijun Zhao
ICASSP6
2025 MUPO-Net: A Multilevel Dual-domain Progressive Enhancement Network with Embedded Attention for CT Metal Artifact Reduction
abstract
Metal implants in patients cause severe streaking artifacts in computed tomography (CT) images, significantly compromising image quality. Deep learning methods have been successfully applied to metal artifact reduction (MAR) in CT, but often result in overly smooth images, failing to reconstruct complex details accurately. In this paper, we propose a multilevel dual-domain progressive enhancement network with embedded attention for MAR, termed MUPO-Net. Our approach constructs a Contrast Weight Mapping (CWM) module that generates a weighted heatmap, allocating weights to different regions based on the influence of metal artifacts, and an ASR-Net (Attention-Embedded Sinogram Restoration Network) that utilizes these weights to better remove artifacts in sinogram domain. Additionally, an Image Detail Enhancement Network (IDE-Net) is proposed to restore fine texture details in CT images through multi-scale feature fusion. Extensive experiments on both synthetic and clinical datasets demonstrate the superior effectiveness of MUPO-Net compared to the state-of-the-art MAR techniques.
Xiaoli Yao, Jia Tan, Zijian Deng, Deng Xiong, Qijun Zhao
ICASSP5
2025 Granularity-Aware Contrastive Learning for Fine-Grained Action Recognition
abstract
The contrastive learning paradigm has been widely used for image-language pre-training and extended to video-text tuning. These approaches aim to maximize the similarity between positive sample pairs while minimizing that of negative ones through an alignment objective. Their performance is highly affected by the definition of positive and negative pairs which depends on the granularity of label classification. This effect is particularly apparent in video action recognition, where different fine-grained actions may belong to a shared coarse label. Therefore, indiscriminately treating a video sample and labels that are not identical at the fine-grained level but share the same coarse label as negative pairs leads to pushing the sample apart from the cluster of its basic coarse action. Such conflict can potentially prevent the model from pulling the sample and its target label closer. For a balanced understanding of coarse and fine-grained distinctions, we propose the Granularity-Aware Contrastive Learning (GACon) framework to improve contrastive learning for fine-grained action recognition. This is achieved through (i) a refined definition of sample-label relation and alignment objectives, and (ii) the exchange of coarse and fine-grained information between two granularity-distinct experts. Experiments on four benchmarks of fine-grained action recognition show the superiority of our proposed GACon compared to existing approaches.
Qijun Zhao
ICASSP3
2025 Identity-Agnostic Learning for Deepfake Face Detection
abstract
Despite the promising results obtained by existing deepfake face detection methods for within-dataset detection, they often fail to generalize effectively to new datasets. We hypothesize that identity, a significant feature in facial recognition, is a key factor affecting deepfake detection models’ cross-dataset performance. In the feature space learned by a real/fake classifier, facial features may cluster based on identity rather than their authenticity, which undermines the classifier’s ability to distinguish between real and fake images. This paper introduces a novel training approach called Identity-Agnostic Learning (IAL) for deepfake face detection. IAL trains the detection model with identity-agnostic manner. It thus guides model to pay attention to the identity-irrelevant features. Experimental results demonstrate that our method effectively enhances the overall generalizability of deepfake face detection models.
Zongyong Deng, Qijun Zhao
ICASSP3
2025 DAPL: Integration of Positive and Negative Descriptions in Text-Based Person Search
abstract
Text-based person search (TBPS) aims to retrieve specific images of individuals from large datasets using textual descriptions. Existing TBPS methods focus primarily on identifying explicit positive attributes, often neglecting the critical role of negative descriptions. This oversight can lead to false positives, where images that should be excluded based on negative descriptions are incorrectly included, due to partial alignment with the positive criteria. To address this limitation, we propose the Dual Attribute Prompt Learning (DAPL) framework, which incorporates both positive and negative descriptions to improve the interpretative accuracy of vision-language models in TBPS tasks. DAPL combines Dual Image-Attribute Contrastive (DIAC) learning with Sensitive Image-Attribute Matching (SIAM) learning to enhance the detection of previously unseen attributes. Furthermore, to achieve a balance between coarse and fine-grained alignment of visual and textual embeddings, we introduce the Dynamic Token-wise Similarity (DTS) loss. This loss function refines the representation of both matching and non-matching descriptions at the token level, providing more precise and adaptable similarity assessments, and ultimately improving the accuracy of the matching process. Empirical results demonstrate that DAPL outperforms state-of-the-art methods, enhancing both precision and robustness in TBPS tasks.
Yuchuan Deng, Zhanpeng Hu, Zijie Xin, Chuang Deng, Qijun Zhao
ICME5
2025 QEMesh: Employing A Quadric Error Metrics-Based Representation for Mesh Generation
abstract
Mesh generation plays a crucial role in 3D content creation, as mesh is widely used in various industrial applications. Recent works have achieved impressive results but still face several issues, such as unrealistic patterns or pits on surfaces, thin parts missing, and incomplete structures. Most of these problems stem from the choice of shape representation or the capabilities of the generative network. To alleviate these, we extend PoNQ, a Quadric Error Metrics (QEM)-based representation, and propose a novel model, QEMesh, for high-quality mesh generation. PoNQ divides the shape surface into tiny patches, each represented by a point with its normal and QEM matrix, which preserves fine local geometry information. In our QEMesh, we regard these elements as generable parameters and design a unique latent diffusion model containing a novel multi-decoder VAE for PoNQ parameters generation. Given the latent code generated by the diffusion model, three parameter decoders produce several PoNQ parameters within each voxel cell, and an occupancy decoder predicts which voxel cells containing parameters to form the final shape. Extensive evaluations demonstrate that our method generates results with watertight surfaces and is comparable to state-of-the-art methods in several main metrics.
Ruowei Wang, Qijun Zhao
ICME4
2025 Promoting Segment Anything Model towards Highly Accurate Dichotomous Image Segmentation
abstract
The Segment Anything Model (SAM) represents a significant breakthrough into foundation models for computer vision, providing a large-scale image segmentation model. However, despite SAM’s zero-shot performance, its segmentation masks lack fine-grained details, particularly in accurately delineating object boundaries. Therefore, it is both interesting and valuable to explore whether SAM can be improved towards highly accurate object segmentation, which is known as the dichotomous image segmentation (DIS) task. To address this issue, we propose DIS-SAM, which advances SAM towards DIS with extremely accurate details. DIS-SAM is a framework specifically tailored for highly accurate segmentation, maintaining SAM’s promptable design. DIS-SAM employs a two-stage approach, integrating SAM with a modified advanced network that was previously designed to handle the prompt-free DIS task. To better train DIS-SAM, we employ a ground truth enrichment strategy by modifying original mask annotations. Despite its simplicity, DIS-SAM significantly advances the SAM, HQ-SAM, and Pi-SAM by ~8.5%, ~6.9%, and ~3.7% maximum F-measure. Our code at https://github.com/Tennine2077/DIS-SAM.
Xianjie Liu, Keren Fu, Yao Jiang 0002, Qijun Zhao
ICME4
2025 Leveraging Hierarchical Spatio-Temporal Distribution Prompt for Zero-Shot Species Recognition
abstract
Reliable species recognition is essential for biodiversity conservation and management. Classic vision-based methods require large amounts of labeled data and suffer from limited generalization capabilities. In contrast, Vision Language Models (VLMs) like CLIP offer a promising alternative by aligning visual representations with text embeddings, enabling effective species recognition even in a zero-shot setting. However, previous works primarily focus on generating species shape and appearance descriptions, which are unreliable for distinguishing visually similar species. Building on the fact that each species has unique temporal and spatial distribution patterns, we introduce a novel prompt called Hierarchical Spatio-Temporal Distribution Prompt to boost CLIP’s zero-shot species recognition performance. Specifically, we construct a hierarchical spatio-temporal distribution prompt using multi-level spatial and temporal distribution descriptions (e.g., from hemisphere to continent and down to country level), integrated with the default CLIP prompt. In this way, text embeddings are enriched with spatio-temporal contextual information, allowing for better alignment with visual features. Extensive experiments conducted on multiple species in the iNaturalist 2021 dataset validate the effectiveness of our proposed method.
Yue Yang 0026, Qijun Zhao
ICME4
2025 Perspective Makes Perfect: Prompt-tuning Vision-Language Models for Action Recognition with Diversified Multi-Modal Observation
abstract
Observing key visual cues and reasoning multi-modal semantics from diversified perspectives are critical in human beings’ perception of actions. However, existing action recognition methods, including ones based on prompt-tuning image-based vision-language (I-VL) models, suffer from the under-exploration of video contents and textual semantics. In this paper, inspired by how human beings reason actions from diversified perspectives, we propose Diversified Multi-Modal Observation (DMMO) to improve prompt-tuning for video action recognition by reducing redundancy and enhancing semantics diversity in various perspectives. Firstly, we select a few informative frames instead of dense sampling for clear visual semantics. Then, subject-context segmentation is applied as diversified visual emphases for textual semantics analysis. Afterward, auxiliary captions with diversified perspectives are generated and aggregated. Finally, we perform multi-modal interaction and align outputs. Extensive evaluation experiments show the superiority of our proposed method, and demonstrate the effectiveness of diversifying prompts from multiple perspectives in tuning I-VL models.
Qijun Zhao, Zhen Zhai
ICME2
2025 Vocalization-based Fine-grained Prediction of Giant Panda Ovulation Time
abstract
The giant panda (Ailuropoda melanoleuca) is a unique animal endemic to China. Predicting the ovulation time of giant pandas is beneficial for better natural mating and artificial insemination decision-making to assist reproduction effectively. Given the shortcomings and uncertainties of biological methods, a Squeeze-and-Excitation ResNet (SE-ResNet)-based model is designed to predict the ovulation time automatically based on the vocalization of giant pandas in breeding seasons. Firstly, we collect and process a dataset of calls from 5 individuals, totaling 7200 seconds. Then, the Mel Spectrogram of the vocalizations is utilized as input for the SE-ResNet and we use balanced sampling to reduce the impact of data imbalance. In addition, we propose a triple Mel filter bank for better feature extraction and perform multi-band feature fusion based on dynamic feature weights. Our experimental results show that the Root Mean Square Error (RMSE) of the proposed method is 4h, which is lower than that of the counterpart baseline methods, demonstrating that the fine-grained prediction of giant panda ovulation time using SE-ResNet based on their vocalizations is feasible and effective.
Jinlong Tang, Qijun Zhao, Rong Hou, Mengnan He
IJCNN2
2025 Camouflaged Object Tracking: A Benchmark
abstract
Visual tracking has seen remarkable advancements, largely driven by the availability of large-scale training datasets that have enabled the development of highly accurate and robust algorithms. While significant progress has been made in tracking general objects, research on more challenging scenarios, such as tracking camouflaged objects, remains limited. Camouflaged objects, which blend seamlessly with their surroundings or other objects, present unique challenges for detection and tracking in complex environments. In critical fields like military, security, agriculture, and marine monitoring, accurately tracking camouflaged objects is essential. To address this gap, we introduce the Camouflaged Object Tracking Dataset (COTD), a specialized benchmark designed specifically for evaluating camouflaged object tracking methods. The COTD dataset comprises 200 sequences and approximately 80,000 frames, each annotated with detailed bounding boxes. Our evaluation of 20 existing tracking algorithms reveals significant deficiencies in their performance with camouflaged objects. To address these issues, we propose a novel tracking framework, HIPTrack-MLS, which demonstrates promising results in improving tracking performance for camouflaged objects. COTD and code are avialable at https://github.com/openat25/HIPTrack-MLS.
Pengzhi Zhong, Defeng Huang, Huikai Shao, Qijun Zhao, Shuiwang Li
ACM Multimedia6
2025 Hierarchical RAG-Driven Multi-hop Reasoning for Medical Video Question Answering
Ruohan Gao, Qijun Zhao, Yangqianqian Chen
NLPCC (4)2
2025 Divide-and-Specialize CLIP: A Text-Guided Multi-expert Framework for Fine-Grained Action Recognition
Lingjie Zeng, Zhen Zhai, Qijun Zhao, Hanyang Lin
PRCV (7)5
2025 Fusing Contextual Clustering with State Space Models for Dzi Beads Image Classification
abstract
The dzi beads pattern categories are rich and diverse, and the differences between the categories are small, which leads to a certain challenge in its classification. To address the problems of weak generalisation performance, low accuracy, and high computational complexity, this paper proposes a FasterVim method, which takes FasterNet as the baseline and uses the designed FasterVimBlock module to improve the overall generalisation performance of the dzi beads images. Firstly, FasterNet is used as a baseline for data enhancement of dzi beads images to improve the overall generalisation performance of the classification model. Secondly, contextual clustering and state space models are fused to effectively implement the interaction between local features of dzi images and the linear representation of image features in long sequences, in order to achieve a balance between the computational complexity and accuracy of the model. Next, point convolution is introduced to achieve different scales of channel attention to efficiently extract the global and local features of the dzi beads image and improve the accuracy of the model. Finally, experimental validation is carried out on the dzi beads image classification dataset, and the proposed method achieves an accuracy of 93.8%, which is 1.91% and 1.5 times higher than the baseline model in Accuracy and Flops, respectively. The experimental results show that the proposed method is competitive in terms of accuracy and number of computational parameters, and can effectively meet the deployment application of the dzi beads image classification model.
Jianjun Xia, Dingguo Gao, Songtao Xu, Qijun Zhao
SMC4
2025 A geometry-aware generative model for face morphing attacks
Zongyong Deng, Qijun Zhao, Libin Ye, Qiaoyun He, Zuyuan He
Knowl. Based Syst.2
2025 Dynamic patch-aware enrichment transformer for occluded person re-identification
Xin Zhang 0125, Keren Fu, Qijun Zhao
Knowl. Based Syst.3
2025 Generalized category discovery in fine-grained recognition
Yifu Yuan, Qijun Zhao
Mach. Vis. Appl.2
2025 Geometric self-supervision for monocular 3D animal pose estimation
Xiaowei Dai, Shuiwang Li, Qijun Zhao, Hongyu Yang 0002
Pattern Recognit.3
2025 Adaptively bypassing vision transformer blocks for efficient visual tracking
Xiangyang Yang 0001, Dan Zeng 0002, Xucheng Wang, You Wu 0009, Hengzhou Ye, Qijun Zhao, Shuiwang Li
Pattern Recognit.6
2025 Exploiting rank-based filter pruning for real-time UAV tracking
Xucheng Wang, Dan Zeng 0002, Qijun Zhao, Shuiwang Li
Signal Process. Image Commun.3
2025 Explicit Motion Handling and Interactive Prompting for Video Camouflaged Object Detection
abstract
Camouflage poses notable challenges in distinguishing a static target, as it usually blends seamlessly with the background. However, any movement by the target can disrupt this disguise, making it detectable. Existing video camouflaged object detection (VCOD) approaches take noisy motion estimation as input or model motion implicitly, restricting detection performance in complex dynamic scenes. In this paper, we propose a novel Explicit Motion handling and Interactive Prompting framework for VCOD, dubbed EMIP, which handles motion cues explicitly using a frozen pre-trained optical flow fundamental model. EMIP is characterized by a two-stream architecture for simultaneously conducting camouflaged segmentation and optical flow estimation. Interactions across the dual streams are realized in an interactive prompting way that is inspired by emerging visual prompt learning. Two learnable modules, i.e. the camouflaged feeder and motion collector, are designed to incorporate segmentation-to-motion and motion-to-segmentation prompts, respectively, and enhance outputs of the both streams. The prompt fed to the motion stream is learned by supervising optical flow in a self-supervised manner. Furthermore, we show that long-term historical information can also be incorporated as a prompt into EMIP and achieve more robust results with temporal consistency. By leveraging promoting techniques based on EMIP, the proposed long-term model EMIP ${}^{\dagger }$ incurs lower training cost with only 8.5M trainable parameters (less than 8% of the total model parameters). Experimental results demonstrate that both EMIP and EMIP ${}^{\dagger }$ set new state-of-the-art records on popular VCOD benchmarks. Additionally, comparative evaluations against other video segmentation models on a wider range of video segmentation tasks demonstrate the robustness and superior generalization capabilities of EMIP. Our code is made publicly available at https://github.com/zhangxin06/EMIP.
Xin Zhang 0125, Ge-Peng Ji, Keren Fu, Qijun Zhao
IEEE Trans. Image Process.6
2025 Imaginarium: Vision-guided High-Quality 3D Scene Layout Generation
abstract
Generating artistic and coherent 3D scene layouts is crucial in digital content creation. Traditional optimization-based methods are often constrained by cumbersome manual rules, while deep generative models face challenges in producing content with richness and diversity. Furthermore, approaches that utilize large language models frequently lack robustness and fail to accurately capture complex spatial relationships. To address these challenges, this paper presents a novel vision-guided 3D layout generation system. We first construct a high-quality asset library containing 2,037 scene assets and 147 3D scene layouts. Subsequently, we employ an image generation model to expand prompt representations into images, fine-tuning it to align with our asset library. We then develop a robust image parsing module to recover the 3D layout of scenes based on visual semantics and geometric information. Finally, we optimize the scene layout using scene graphs and overall visual semantics to ensure logical coherence and alignment with the images. Extensive user testing demonstrates that our algorithm significantly outperforms existing methods in terms of layout richness and quality. The code and dataset will be available at https://github.com/HiHiAllen/Imaginarium.
Qinghongbing Xie, Junsheng Yu, Yirui Guan, Zhongyuan Liu, Qijun Zhao, Ligang Liu 0001, Long Zeng 0001
ACM Trans. Graph.9
2024 Hierarchical Generative Network for Face Morphing Attacks
abstract
Face morphing attacks circumvent face recognition systems (FRSs) by creating a morphed image that contains multiple identities. However, existing face morphing attack methods either sacrifice image quality or compromise the identity preservation capability. Consequently, these attacks fail to bypass FRSs verification well while still managing to deceive human observers. These methods typically rely on global information from contributing images, ignoring the detailed information from effective facial regions. To address the above issues, we propose a novel morphing attack method to improve the quality of morphed images and better preserve the contributing identities. Our proposed method leverages the hierarchical generative network to capture both local detailed and global consistency information. Additionally, a mask-guided image blending module is dedicated to removing artifacts from areas outside the face to improve the image's visual quality. The proposed attack method is compared to state-of-the-art methods on three public datasets in terms of FRSs' vulnerability, attack detectability, and image quality. The results show our method's potential threat of deceiving FRSs while being capable of passing multiple morphing attack detection (MAD) scenarios.
Zuyuan He, Zongyong Deng, Qiaoyun He, Qijun Zhao
FG4
2024 Text and Edge Guided Thangka Image Inpainting with Diffusion Model
abstract
Thangka images, revered as the Tibetan encyclopedia for their diverse content, represent an invaluable cultural heritage. However, years of display, veneration, and improper preservation often led to varying degrees of damage to Thangka paintings. High-quality Thangka image inpainting, where damaged regions are filled with plausible content according to prior information, remains a significant challenge. Despite the excellent performance of recent text-to-image diffusion models have attracted extensive attention due to their excellent performance in generating various high-quality natural images. However, they encounter challenges in understanding textual descriptions related to complex Thangka images. To address this, we propose a new two-stage inpainting framework that employs both text and edge guidance. First, we employ GAN-based edge prediction network to predict the missing edges. Subsequently, these predicted edges guide the diffusion model through ControlNet, enhancing the inpainting areas and blending them with the original image. Additionally, we incorporate multiple LoRA models trained on Thangka datasets to integrate text guidance with edge control. Experiments results showed the effectiveness of our method in inpainting Thangka images with plausible visual consistency and preserved details.
Tienyi Hsieh, Qijun Zhao, Pubu Danzeng, Dingguo Gao, Dorji Gesang
ICME2
2024 Face Morphing via Adversarial Attack-based Adaptive Blending
abstract
In this paper, we propose an innovative adversarial attack-based adaptive blending architecture (A3B) for face morphing attacks. Unlike traditional face morphing methods that evenly blend identity information using a half-half-hard strategy, we propose an approach that considers the varying importance of facial features in different regions of individuals for more precise and clear face morphing. Our method adaptively blends the latent codes of contributing subjects with a weighting mask that assigns different weights to different facial features when blending two faces to fuse their identities. The crux of our method lies in computing the weighting mask given a pair of contributing face images. This is done with the assistance of adversarial attacks, which can adaptively perturb face images to conceal or transfer identity. Owing to the adaptive blending strategy, our proposed approach achieves competitive performance on several state-of-the-art benchmark datasets. In contrast to existing methods, our approach explores the connection between adversarial attacks and morphing attacks for the first time, which generates morphed face images with plausible visual quality and simultaneously preserves the identity of contributing subjects. This novel perspective raises concerns about the potential security risks it poses to current facial recognition systems.
Qiaoyun He, Zongyong Deng, Zuyuan He, Qijun Zhao
IJCNN4
2024 Feature-Aware Noise Contrastive Learning For Unsupervised Red Panda Re-Identification
abstract
To facilitate the re-identification (re-ID) of individual animals, existing methods primarily focus on maximizing feature similarity within the same individual and enhancing distinctiveness between different individuals. However, most of them still rely on supervised learning and require substantial labeled data, which is challenging to obtain. To avoid this issue, we propose a Feature-Aware Noise Contrastive Learning (FANCL) method to explore an unsupervised learning solution, which is then validated on the task of red panda re-ID. FANCL employs a Feature-Aware Noise Addition module to produce noised images that conceal critical features, and designs two contrastive learning modules to calculate the losses. Firstly, a feature consistency module is designed to bridge the gap between the original and noised features. Secondly, the neural networks are trained through a cluster contrastive learning module. Through these more challenging learning tasks, FANCL can adaptively extract deeper representations of red pandas. The experimental results on a set of red panda images collected in both indoor and outdoor environments prove that FANCL outperforms several related state-of-the-art unsupervised methods, achieving high performance comparable to supervised learning methods.
Qijun Zhao
IJCNN2
2024 Towards Labeling-free Fine-grained Animal Pose Estimation
Dan Zeng 0002, Shuiwang Li, Qijun Zhao, Qiaomu Shen, Bo Tang 0016
ACM Multimedia4
2024 GenUDC: High Quality 3D Mesh Generation With Unsigned Dual Contouring Representation
Ruowei Wang, Dan Zeng 0002, Xueqi Ma, Zixiang Xu, Jianwei Zhang 0013, Qijun Zhao
ACM Multimedia7
2024 MTFusion: Reconstructing Any 3D Object from Single Image Using Multi-word Textual Inversion
Ruowei Wang, Zixiang Xu, Qijun Zhao
PRCV (6)5
2024 Identity-Preserving Animal Image Generation for Animal Individual Identification
Zongming Peng, Yangqianqian Chen, Keren Fu, Qijun Zhao
PRCV (15)7
2024 Key Object Detection: Unifying Salient and Camouflaged Object Detection Into One Task
Pengyu Yin, Keren Fu, Qijun Zhao
PRCV (12)3
2024 Species-Aware Guidance for Animal Action Recognition with Vision-Language Knowledge
Zhen Zhai, Qijun Zhao, Keren Fu
PRCV (7)3
2024 Prior Mask-Guided Highly Accurate Dichotomous Image Segmentation
Shanfeng Zhou, Keren Fu, Qijun Zhao
PRICAI (4)5
2024 LSTPNet: Long short-term perception network for dynamic facial expression recognition in the wild
Chengcheng Lu, Yiben Jiang, Keren Fu, Qijun Zhao, Hongyu Yang 0002
Image Vis. Comput.4
2024 Polarimetric inverse scattering using LSTM-aided association-learning sparse reconstruction framework
Yue Yang 0026, Qijun Zhao, Qun Wan
Signal Process.2
2024 Parallax-Aware Network for Light Field Salient Object Detection
abstract
Multi-view images capture scene details from different views, making them advantageous for light field salient object detection (LF SOD). However, most existing LF SOD methods neglect effective modeling and utilization of parallax information inherent in multi-view images. To address this limitation, we propose to explicitly model parallax information and conduct SOD in a parallax-aware manner, resulting in a novel network called PANet. Our model initiates by generating horizontal and vertical visual parallax maps from four border views using optical flow estimation. We then introduce a parallax-aware network, incorporating a parallax processing module (PPM) that handles both parallax quality assessment and parallax correction. In the parallax correction phase, we design a channel-based correction unit (CCU) and a graph-based correction unit (GCU) to rectify deviations of parallax features in a direction-specific manner. Additionally, a parallax supplement module (PSM) seamlessly fuses the parallax information from different directions and embeds it into the center view, thereby improving SOD accuracy. Experiments on three benchmark datasets demonstrate the superiority of our PANet model over 15 state-of-the-art models. Our code for the model will be publicly available soon.
Yao Jiang 0002, Keren Fu, Qijun Zhao
IEEE Signal Process. Lett.4
2024 Learning Target-Aware Vision Transformers for Real-Time UAV Tracking
abstract
In recent years, the field of unmanned aerial vehicle (UAV) tracking has grown rapidly, finding numerous applications across various industries. While the discriminative correlation filters (DCF)-based trackers remain the most efficient and widely used in the UAV tracking, recently lightweight convolutional neural network (CNN)-based trackers using filter pruning have also demonstrated impressive efficiency and precision. However, the performance of these lightweight CNN-based trackers is still far from satisfactory. In the generic visual tracking, emerging vision transformer (ViT)-based trackers have shown great success by using cross-attention instead of correlation operation, enabling more effective capturing of relationships between the target and the search image. But to best of the authors’ knowledge, the UAV tracking community has not yet well explored the potential of ViTs for more effective and efficient template-search coupling for UAV tracking. In this article, we propose an efficient ViT-based tracking framework for real-time UAV tracking. Our framework integrates feature learning and template-search coupling into an efficient one-stream ViT to avoid an extra heavy relation modeling module. However, we observe that it tends to weaken the target information through transformer blocks due to the significantly more background tokens. To address this problem, we propose to maximize the mutual information (MI) between the template image and its feature representation produced by the ViT. The proposed method is dubbed TATrack. In addition, to further enhance efficiency, we introduce a novel MI maximization-based knowledge distillation, which strikes a better trade-off between accuracy and efficiency. Exhaustive experiments on five benchmarks show that the proposed tracker achieves state-of-the-art performance in UAV tracking. Code is released at:https://github.com/xyyang317/TATrack.
Shuiwang Li, Xiangyang Yang 0001, Xucheng Wang, Dan Zeng 0002, Hengzhou Ye, Qijun Zhao
IEEE Trans. Geosci. Remote. Sens.6
2024 Transformer-Based Light Field Salient Object Detection and Its Application to Autofocus
abstract
Existing light field salient object detection (LFSOD) models predominantly rely on convolutional neural networks or local attention to process light field data, consequently encountering difficulties in modeling intra-slice and cross-slice long-range dependencies within focal stacks. In this paper, we ponder the feasibility of relying solely on the pure Transformer architecture to address this dilemma and propose a novel quasi-pure Transformer-based framework for LFSOD, termed TLFNet. TLFNet incorporates innovative Transformer-based fusion modules (PGFormer) along with an edge enhancement module. The PGFormer employs a perpendicular self-attention (PSA) mechanism to capture long-range dependencies along both cross-slice and intra-slice axes within the focal stack, and integrates multi-modal features using a guided feature fusion (GFF) module. To address the issue of blurry edges arising from the Transformer-based encoder-decoder architecture, the edge enhancement module combines detailed texture and body information and employs focal loss to improve the edge precision of salient objects. TLFNet is a nearly pure Transformer-based approach (with approximately 99.01% of its parameters belonging to the Transformer), while the edge enhancement module significantly boosts accuracy with only around 0.99% of parameters. Comprehensive benchmarks demonstrate that TLFNet outperforms 14 light field models and achieves new state-of-the-art performance. Last but not least, we show in this paper a new application scheme of TLFNet, by cooperating with the deep autofocus technique proposed in [1], leading to light field salient object autofocus (LFSOA). LFSOA aims to identify and output the focal slice with a salient object in focus while keeping other irrelevant background blurred (out-of-focus), yielding an autonomous bokeh effect in photography. The code for the model and application will be publicly available soon.
Yao Jiang 0002, Keren Fu, Qijun Zhao
IEEE Trans. Image Process.4
2024 Salient Object Detection in RGB-D Videos
abstract
Given the widespread adoption of depth-sensing acquisition devices, RGB-D videos and related data/media have gained considerable traction in various aspects of daily life. Consequently, conducting salient object detection (SOD) in RGB-D videos presents a highly promising and evolving avenue. Despite the potential of this area, SOD in RGB-D videos remains somewhat under-explored, with RGB-D SOD and video SOD (VSOD) traditionally studied in isolation. To explore this emerging field, this paper makes two primary contributions: the dataset and the model. On one front, we construct the RDVS dataset, a new RGB-D VSOD dataset with realistic depth and characterized by its diversity of scenes and rigorous frame-by-frame annotations. We validate the dataset through comprehensive attribute and object-oriented analyses, and provide training and testing splits. Moreover, we introduce DCTNet+, a three-stream network tailored for RGB-D VSOD, with an emphasis on RGB modality and treats depth and optical flow as auxiliary modalities. In pursuit of effective feature enhancement, refinement, and fusion for precise final prediction, we propose two modules: the multi-modal attention module (MAM) and the refinement fusion module (RFM). To enhance interaction and fusion within RFM, we design a universal interaction module (UIM) and then integrate holistic multi-modal attentive paths (HMAPs) for refining multi-modal low-level features before reaching RFMs. Comprehensive experiments, conducted on pseudo RGB-D video datasets alongside our proposed RDVS, highlight the superiority of DCTNet+ over 19 VSOD models and 14 RGB-D SOD models. Additionally, insightful ablation experiments were performed on both pseudo and realistic RGB-D video datasets to demonstrate the advantages of individual modules as well as the necessity of introducing realistic depth into VSOD. Our code together with RDVS dataset will be available at https://github.com/kerenfu/RDVS/.
Ao Mou, Yukang Lu, Dingyao Min, Keren Fu, Qijun Zhao
IEEE Trans. Image Process.6
2024 3-D Convolutional Neural Networks for RGB-D Salient Object Detection and Beyond
abstract
RGB-depth (RGB-D) salient object detection (SOD) recently has attracted increasing research interest, and many deep learning methods based on encoder-decoder architectures have emerged. However, most existing RGB-D SOD models conduct explicit and controllable cross-modal feature fusion either in the single encoder or decoder stage, which hardly guarantees sufficient cross-modal fusion ability. To this end, we make the first attempt in addressing RGB-D SOD through 3-D convolutional neural networks. The proposed model, named RD3D, aims at prefusion in the encoder stage and in-depth fusion in the decoder stage to effectively promote the full integration of RGB and depth streams. Specifically, RD3D first conducts prefusion across RGB and depth modalities through a 3-D encoder obtained by inflating 2-D ResNet and later provides in-depth feature fusion by designing a 3-D decoder equipped with rich back-projection paths (RBPPs) for leveraging the extensive aggregation ability of 3-D convolutions. Toward an improved model RD3D+, we propose to disentangle the conventional 3-D convolution into successive spatial and temporal convolutions and, meanwhile, discard unnecessary zero padding. This eventually results in a 2-D convolutional equivalence that facilitates optimization and reduces parameters and computation costs. Thanks to such a progressive-fusion strategy involving both the encoder and the decoder, effective and thorough interactions between the two modalities can be exploited and boost detection accuracy. As an additional boost, we also introduce channel-modality attention and its variant after each path of RBPP to attend to important features. Extensive experiments on seven widely used benchmark datasets demonstrate that RD3D and RD3D+ perform favorably against 14 state-of-the-art RGB-D SOD approaches in terms of five key evaluation metrics. Our code will be made publicly available at https://github.com/PPOLYpubki/RD3D.
Yanye Lu, Keren Fu, Qijun Zhao
IEEE Trans. Neural Networks Learn. Syst.5
2024 Event-Triggered Fractional-Order Tracking Control for an Uncertain Nonlinear System With Output Saturation and Disturbances
abstract
In this article, an event-triggered (ET) fractional-order adaptive tracking control scheme (ATCS) is studied for the uncertain nonlinear system with the output saturation and the external disturbances by using the nonlinear disturbance observer (NDO) and the neural networks (NNs). Based on NNs, the system uncertainties are approximated. An NN-based NDO is designed to estimate the bounded disturbances. Combining the NNs, the output of the designed NDO, the fractional-order theory, and the ET mechanism, an ATCS is proposed under the output saturation. According to the stability analysis, all the closed-loop signals are semiglobally uniformly ultimately bounded based on the investigative ATCS. The simulation results and the comparative experiment verifications are shown to indicate the viability of the developed control scheme.
Shuyi Shao, Mou Chen, Sijia Zheng, Shumin Lu, Qijun Zhao
IEEE Trans. Neural Networks Learn. Syst.5
2023 Unsupervised 3D Animal Canonical Pose Estimation with Geometric Self-Supervision
abstract
Although analyzing animal shape and pose has potential applications in many fields, there is little work on 3D animal pose estimation. This can be attributed to two aspects: the lack of large-scale well-annotated datasets, and perspective ambiguities which make it difficult to map 2D space to 3D space. To address data scarcity, we propose an unsupervised method to estimate 3D animal pose, given only 2D poses. To deal with perspective ambiguities, we introduce a canonical consistency loss and a camera consistency loss to impose geometric priors in the training process, and combine the reprojection loss and the 2D pose discriminator to enable self-supervised learning. Specifically, given a 2D pose, the pose generator network generates a corresponding 3D pose and the camera network estimates a camera rotation. During training, the generated 3D pose is randomly reprojected onto camera viewpoints to synthesize a new 2D pose. The synthesized 2D pose is decomposed into a 3D pose and a camera rotation, based on which consistency losses are imposed in both 3D canonical poses and camera rotations for self-supervised training. We evaluate the proposed method on real and synthetic datasets, i.e., SMAL and AcinoSet. The experimental results demonstrate the effectiveness of the proposed method and we achieve state-of-the-art performance among unsupervised algorithms for 3D animal canonical pose estimation.
Xiaowei Dai, Shuiwang Li, Qijun Zhao, Hongyu Yang 0002
FG3
2023 Cascaded Network-Based Single-View Bird 3D Reconstruction
Pei Su, Qijun Zhao
ICANN (2)2
2023 Optimal-Landmark-Guided Image Blending for Face Morphing Attacks
abstract
In this paper, we propose a novel approach for conducting face morphing attacks, which utilizes optimal-landmark-guided image blending. Current face morphing attack can be categorized into landmark-based and generation-based approaches. Landmark-based methods use geometric transformations to warp facial regions according to averaged landmarks, but often produce morphed images with poor visual quality. Generation-based methods, which employ generation models to blend multiple face images, can achieve better visual quality, but are often unsuccessful in generating morphed images that can effectively evade state-of-the-art face recognition systems (FRSs). Our proposed method overcomes the limitations of previous approaches by optimizing the morphing landmarks and using Graph Convolutional Networks (GCNs) to combine landmark and appearance features. We model facial landmarks as nodes in a bipartite graph that is fully connected, and utilize GCNs to simulate their spatial and structural relationships. The aim is to capture variations in facial shape and enable accurate manipulation of facial appearance features during the warping process, resulting in morphed facial images that are highly realistic and visually faithful. Experiments on two public datasets prove that our method inherits the advantages of previous landmark-based and generation-based methods and generates morphed images with higher quality, posing a more significant threat to state-of-the-art FRSs.
Qiaoyun He, Zongyong Deng, Zuyuan He, Qijun Zhao
IJCB4
2023 3D Semantic Subspace Traverser: Empowering 3D Generative Model with Shape Editing Capability
abstract
Shape generation is the practice of producing 3D shapes as various representations for 3D content creation. Previous studies on 3D shape generation have focused on shape quality and structure, without or less considering the importance of semantic information. Consequently, such generative models often fail to preserve the semantic consistency of shape structure or enable manipulation of the semantic attributes of shapes during generation. In this paper, we proposed a novel semantic generative model named 3D Semantic Subspace Traverser that utilizes semantic attributes for category-specific 3D shape generation and editing. Our method utilizes implicit functions as the 3D shape representation and combines a novel latent-space GAN with a linear subspace model to discover semantic dimensions in the local latent space of 3D shapes. Each dimension of the subspace corresponds to a particular semantic attribute, and we can edit the attributes of generated shapes by traversing the coefficients of those dimensions. Experimental results demonstrate that our method can produce plausible shapes with complex structures and enable the editing of semantic attributes. The code and trained models are available at https://github.com/TrepangCat/3D Semantic Subspace Tra verser
Ruowei Wang, Pei Su, Jianwei Zhang 0013, Qijun Zhao
ICCV5
2023 Guided Focal Stack Refinement Network for Light Field Salient Object Detection
abstract
Light field salient object detection (SOD) is an emerging research direction attributed to the richness of light field data. However, most existing methods lack effective handling of focal stacks, therefore making the latter involved in a lot of interfering information and degrade the performance of SOD. To address this limitation, we propose to utilize multi-modal features to refine focal stacks in a guided manner, resulting in a novel guided focal stack refinement network called GFRNet. To this end, we propose a guided refinement and fusion module (GRFM) to refine focal stacks and aggregate multi-modal features. In GRFM, all-in-focus (AiF) and depth modalities are utilized to refine focal stacks separately, leading to two novel sub-modules for different modalities, namely AiF-based refinement module (ARM) and depth-based refinement module (DRM). Such refinement modules enhance structural and positional information of salient objects in focal stacks, and are able to improve SOD accuracy. Experimental results on four benchmark datasets demonstrate the superiority of our GFRNet model against 12 state-of-the-art models.
Yao Jiang 0002, Keren Fu, Qijun Zhao
ICME4
2023 Counting and Locating Anything: Class-agnostic Few-shot Object Counting and Localization
abstract
We focus on a new task for class-agnostic object localization (CAOL), which not only counts objects of any class but also outputs their point locations. Given one or more examples of objects of interest, the CAOL model should output the count and point locations of objects of interest in the query image. We propose two solutions: (1) a post-processing method for class-agnostic object counting models based on density maps, called Maximum Search Strategy (MSS), which converts density maps into specific point coordinates without the need for a threshold. (2) a point-based Class-Agnostic Counting and Localization (CACL) model that directly predicts points for each object of interest, providing an end-to-end solution without the need for post-processing; While MSS obtains promising localization results as long as good quality density map is available, CACL achieves state-of-the-art results for both counting and localization tasks. Our experiments on the CARPK dataset demonstrate the strong generalization performance of CACL model.
Qijun Zhao
ICME3
2023 ConCAP: Contrastive Context-Aware Prompt for Resource-hungry Action Recognition
abstract
Existing large-scale image-language pre-trained models, e.g., CLIP [1], have revealed strong spatial recognition capability on various vision tasks. However, they achieve inferior performance in action recognition due to lack of temporal reasoning ability. Moreover, fully tuning large models require expensive computational infrastructures, and state-of-the-art video models yield slow inference speed due to the high frame sampling rate. The above drawbacks make existing video action recognition works impractical to be applied in resource-hungry scenarios, which is common in the real world. In this work, we propose Contrastive Context-Aware Prompt (ConCAP) for resource-hungry action recognition. Specifically, we develop a lightweight PromptFormer to learn the spatio-temporal representations stacking on top of frozen frame-wise visual backbones, where learnable prompt tokens are plugged between frame tokens during self-attention. These prompt tokens are expected to auto-complete the contextual spatiotemporal information between frames and therefore enhance the model’s representation capability. To achieve this goal, we align the prompt-enhanced representation with both category-level textual representations and video representations from densely sampled frames. Extensive experiments on four video benchmarks show that we achieve state-of-the-art or competitive performance compared to existing methods with far fewer trainable parameters and faster inference speed with limited frames, demonstrating the superiority of ConCAP in resource-hungry scenarios.
Ziyun Zeng, Qijun Zhao, Zhen Zhai
ICME3
2023 Video-based Red Panda Individual Identification by Adaptively Aggregating Discriminative Features
abstract
Existing animal individual identification methods mostly use only single images and cannot effectively leverage complementary features in video frames. To further improve the robustness and accuracy of individual identification of animals like red pandas that have complex body deformation or pose variations, we propose in this paper a deep network to learn hybrid feature representation of red pandas that adaptively aggregates local and global features for red panda identification. The local feature representation is obtained by adaptively finding discriminative local patches of the red panda in each frame and aggregating the local features across frames via a hypergraph neural network. The global feature representation is obtained by aggregating the features of different frames via average pooling. Red panda individuals are finally identified based on the concatenated local and global feature representations. Evaluation experiments have been done on a self-collected dataset of red panda videos. The results prove the utility of video data in animal individual identification as well as the superiority of our proposed method in exploiting discriminative features in video frames for identifying individual animals.
Mengnan He, Qijun Zhao
IJCNN8
2023 Long Short-Term Perception Network for Dynamic Facial Expression Recognition
Chengcheng Lu, Yiben Jiang, Keren Fu, Qijun Zhao, Hongyu Yang 0002
PRCV (5)4
2023 Non-local Temporal Modeling for Practical Skeleton-Based Gait Recognition
Pengyu Peng, Zongyong Deng, Feiyu Zhu 0001, Qijun Zhao
PRCV (5)4
2023 An adaptive matching control method of multiple turboshaft engines
Yong Wang 0016, Chuang Ji, Zhihua Xi, Haibo Zhang 0004, Qijun Zhao
Eng. Appl. Artif. Intell.5
2023 Automatic identification of individual yaks in in-the-wild images using part-based convolutional networks with self-supervised learning
Cuo Da, Qijun Zhao, Liyuan Zhou, Suonan Jiancuo
Expert Syst. Appl.4
2023 Fair Communications in UAV Networks for Rescue Applications
abstract
We study the deployment of an unmanned aerial vehicle (UAV) network to provide urgent communications to people trapped in a disaster zone, where each UAV is an aerial base station in the air. Unlike most existing studies that assumed that each user communicates with a UAV directly, we introduce Device-to-Device (D2D) communications, in which a user within the communication range of a UAV can serve as a hotspot (e.g., WiFi hotspot), and provide communication services to his nearby users who are out of the communication range of any UAV. More users thus can have the communication service provided by the UAV network. To ensure that the users within and out of the communication ranges of deployed UAVs havefaircommunication quality, we study a novel UAV deployment and resource allocation problem under the D2D communication model, which is to deploy$K$given UAVs in the top of a disaster zone, allocate the bandwidth of each UAV to its served users, allocate the bandwidth of each hotspot to his served users, determine the data rate of each user, and find the routing paths for data transmissions, such that the accumulative utility of all users is maximized. We also propose a novel$(1-1/e-\epsilon)$-approximation algorithmalgMaxUtilityfor the problem, where$e$is the base of the natural logarithm, and$\epsilon $is a given constant with$0 < \epsilon < 1-1/e$. We finally evaluate the performance of the algorithm. Experimental results show that accumulative utility by the algorithm is up to 18% larger than those by existing algorithms. In addition, more than 16% users are served in the deployed UAV network by the proposed algorithm.
Qunli Shen, Jian Peng 0002, Wenzheng Xu, Yueying Sun, Weifa Liang, Liangyin Chen, Qijun Zhao, Xiaohua Jia
IEEE Internet Things J.7
2022 Evaluating the perceived safety of urban city via maximum entropy deep inverse reinforcement learning
Yaxuan Wang, Zhixin Zeng, Qijun Zhao
ACML3
2022 Animal Pose Refinement in 2D Images with 3D Constraints
Xiaowei Dai, Shuiwang Li, Qijun Zhao, Hongyu Yang 0002
BMVC3
2022 Unsupervised Homography Estimation with Coplanarity-Aware GAN
abstract
Estimating homography from an image pair is a fundamental problem in image alignment. Unsupervised learning methods have received increasing attention in this field due to their promising performance and label-free training. However, existing methods do not explicitly consider the problem of plane-induced parallax, which will make the predicted homography compromised on multiple planes. In this work, we propose a novel method HomoGAN to guide unsupervised homography estimation to focus on the dominant plane. First, a multi-scale transformer network is designed to predict homography from the feature pyramids of input images in a coarse-to-fine fashion. Moreover, we propose an unsupervised GAN to impose coplanarity constraint on the predicted homography, which is realized by using a generator to predict a mask of aligned regions, and then a discriminator to check if two masked feature maps are induced by a single homography. To validate the effectiveness of HomoGAN and its components, we conduct extensive experiments on a large-scale dataset, and results show that our matching error is 22% lower than the previous SOTA method. Code is available at https://github.com/megvii-research/HomoGAN
Mingbo Hong, Nianjin Ye, Chunyu Lin, Qijun Zhao, Shuaicheng Liu
CVPR5
2022 Depth-Cooperated Trimodal Network for Video Salient Object Detection
abstract
Depth can provide useful geographical cues for salient object detection (SOD), and has been proven helpful in recent RGB-D SOD methods. However, existing video salient object detection (VSOD) methods only utilize spatiotemporal information and seldom exploit depth information for detection. In this paper, we propose a depth-cooperated trimodal network, called DCTNet for VSOD, which is a pioneering work to incorporate depth information to assist VSOD. To this end, we first generate depth from RGB frames, and then propose an approach to treat the three modalities unequally. Specifically, a multi-modal attention module (MAM) is designed to model multi-modal long-range dependencies between the main modality (RGB) and the two auxiliary modalities (depth, optical flow). We also introduce a refinement fusion module (RFM) to suppress noises in each modality and select useful information dynamically for further feature refinement. Lastly, a progressive fusion strategy is adopted after the refined features to achieve final cross-modal fusion. Experiments on five benchmark datasets demonstrate the superiority of our depth-cooperated model against 12 state-of-the-art methods, and the necessity of depth is also validated.
Yukang Lu, Dingyao Min, Keren Fu, Qijun Zhao
ICIP4
2022 A Dataset and Method for Gait Recognition with Unmanned Aerial Vehicless
abstract
Gait is a promising behavioral biometric trait that can be used to identify persons at a distance. Thanks to its advantages of long working distances, low requirements for data quality, and no requirement for cooperation, gait recognition is quite suitable for deployment with unmanned aerial vehicles (UAVs). However, existing works all consider gait recognition in ground-view videos rather than aerial videos. To fill this gap, this paper for the first time constructs a dataset of both UAV-view and ground-view gait data, namely UAV-Gait, which contains 9,898 gait sequences of 202 individuals captured by one DJI consumer UAV flying at altitudes of 10, 20, and 30 meters as well as by five cameras fixed on the ground. Moreover, this paper proposes a graph convolution based part feature pooling method to improve the robustness of extracted gait features to large view changes, especially to pitch rotations, that are common in UAV-captured images. Evaluation experiments on the UAV-Gait dataset prove that gait recognition with UAVs is very challenging and the proposed method can effectively increase the accuracy of gait recognition in aerial images.
Qijun Zhao, Pengyu Peng
ICME2
2022 Rank-Based Filter Pruning for Real-Time UAV Tracking
abstract
Unmanned aerial vehicle (UAV) tracking has wide poten-tial applications in such as agriculture, navigation, and public security. However, the limitations of computing resources, battery capacity, and maximum load of UAV hinder the de-ployment of deep learning-based tracking algorithms on UAV. Consequently, discriminative correlation filters (DCF) track-ers stand out in the UAV tracking community because of their high efficiency. However, their precision is usually much lower than trackers based on deep learning. Model compression is a promising way to narrow the gap (i.e., effciency, precision) between DCF- and deep learning- based trackers, which has not caught much attention in UAV tracking. In this paper, we propose the P-SiamFC++ tracker, which is the first to use rank-based filter pruning to compress the SiamFC++ model, achieving a remarkable balance between efficiency and precision. Our method is general and may encourage further studies on UAV tracking with model compression. Extensive experiments on four UAV benchmarks, including UAV123@10fps, DTB70, UAVDT and Vistrone2018, show that P-SiamFC++ tracker significantly outperforms state-of-the-art UAV tracking methods.
Xucheng Wang, Dan Zeng 0002, Qijun Zhao, Shuiwang Li
ICME3
2022 Local-Global Interaction and Progressive Aggregation for Video Salient Object Detection
Dingyao Min, Chao Zhang 0072, Yukang Lu, Keren Fu, Qijun Zhao
ICONIP (6)5
2022 Learning Multi-Granularity Features for Re-Identifying Figures in Portrait Thangka Images
abstract
Re-identifying the figures in portrait Thangka images is of great significance to the protection and dissemination of Thangka. Existing portrait Thangka image retrieval methods are limited in extracting features of multiple granularities, and consequently can not distinguish between Thangka figures which have different identities but subtle visual differences. To cope with this challenging task, we propose a Multi-Granularity Feature Learning (MGFL) framework consisting of part and global branches. The part branch splits feature maps along the vertical dimension with inspiration by the composition rules of portrait Thangka images. The global branch integrates inception feature pyramid, channel and spatial attention mechanisms to capture discriminative features of various granularities. To combine the multi-granularity features, we propose to use step-wise graph convolution to better exploit the relationships among different composition elements in portrait Thangka images. Extensive qualitative and quantitative experiments on a dataset with 6,208 portrait Thangka images of 145 identities collected by ourselves demonstrate the superiority of our method. We will release the dataset and code to the public for research purpose.
Xire Danzeng, Qijun Zhao, Pubu Danzeng, Xinsheng Li, Gesang Duoji, Dingguo Gao
ICPR4
2022 Learning Generalisable Representations for Offline Signature Verification
abstract
Current offline signature verification methods based on deep learning have achieved promising results, but these methods degrade greatly in cross-domain settings. An efficient offline signature verification model with both high performance and for deployment cross-domain without any adaptation. In this paper, we propose a novel approach to learning generalisable representations for offline signature verification. Firstly, we use the Siamese network combined with Triplet loss and Cross Entropy (CE) loss to learn discriminative features. Secondly, we introduce Instance Normalization (IN) into the network to cope with cross-domain discrepancies and propose an Inference Layer Normalization Neck (ILNNeck) module to further improve model generalization. We evalute the method on our self-collected Multilingual Signature dataset (MLSig) and three public datasets: BHSig-H, BHSig-B, and CEDAR. Results show that while our method achieves comparable results in single-domain setting, it is obviously superior to state-of-the-art methods in cross-domain setting.
Xianmu Cairang, Duoji Zhaxi, Yan Hou, Qijun Zhao, Dingguo Gao, Pubu Danzeng, Dorji Gesang
IJCNN5
2022 Look longer to see better: Audio-visual event localization by exploiting long-term correlation
abstract
Visual and auditory modalities both contain a large amount of rich information about audio-visual events. While the human perception system can effectively fuse the information of the dual modalities in recognizing events, it is still an open issue how to effectively integrate dual-modal information for the task of automatic localization of audio-visual events in videos. In this paper, we propose an audio-visual long-term correlation network to capture the longer correlation of audio and visual features, which is underused by existing methods. To this end, we first propose the time-spatial guided attention (TSGA) module, which locates the spatial region of the audio-visual events in the video and focuses on continuous changes in that location. We then propose the positive time residual fusion (PTRF) module, which encodes the temporal correlation matrix of video and audio, and uses residual fusion to combine audio and visual features. We finally evaluate our method for the fully supervised and weakly supervised tasks on the AVE dataset. The results prove the superiority of our method over its counterparts.
Longyin Guo, Qijun Zhao, Hongmei Gao
IJCNN2
2022 Towards Causality Inference for Very Important Person Localization
abstract
Very Important Person Localization (VIPLoc) aims at detecting certain individuals in a given image, who are more attractive than others in the image. Existing uncontrolled VIPLoc benchmark assumes that the image has one single VIP, which is not suitable for actual application scenarios when multiple VIPs or no VIPs appear in the image. In this paper, we re-built a complex uncontrolled conditions (CUC) dataset to make the VIPLoc closer to the actual situation, containing no, single, and multiple VIPs. Existing methods use the hand-designed and deep learning strategies to extract the features of persons and analyze the differences between VIPs and other persons from the perspective of statistics. They are not explainable as to why the VIP located this output for that input. Thus, there exist the severe performance degradation when we use these models in real-world VIPLoc. Specifically, we establish a causal inference framework that unpacks the causes of previous methods and derives a new principled solution for VIPLoc. It treats the scene as confounding factor, allowing the ever-elusive confounding effects to be eliminated and the essential determinants to be uncovered. Through extensive experiments, our method outperforms the state-of-the-art methods on public VIPLoc datasets and the re-built CUC dataset.
Xiao Wang 0029, Zheng Wang 0007, Wu Liu 0005, Xin Xu 0007, Qijun Zhao, Shin'ichi Satoh 0001
ACM Multimedia5
2022 Light field salient object detection: A review and benchmark
abstract
Salient object detection (SOD) is a long-standing research topic in computer vision with increasing interest in the past decade. Since light fields record comprehensive information of natural scenes that benefit SOD in a number of ways, using light field inputs to improve saliency detection over conventional RGB inputs is an emerging trend. This paper provides the first comprehensive review and a benchmark for light field SOD, which has long been lacking in the saliency community. Firstly, we introduce light fields, including theory and data forms, and then review existing studies on light field SOD, covering ten traditional models, seven deep learning-based models, a comparative study, and a brief review. Existing datasets for light field SOD are also summarized. Secondly, we benchmark nine representative light field SOD models together with several cutting-edge RGB-D SOD models on four widely used light field datasets, providing insightful discussions and analyses, including a comparison between light field SOD and RGB-D SOD models. Due to the inconsistency of current datasets, we further generate complete data and supplement focal stacks, depth maps, and multi-view images for them, making them consistent and uniform. Our supplemental data make a universal benchmark possible. Lastly, light field SOD is a specialised problem, because of its diverse data representations and high dependency on acquisition hardware, so it differs greatly from other saliency detection tasks. We provide nine observations on challenges and future directions, and outline several open issues. All the materials including models, datasets, benchmarking results, and supplemented light field datasets are publicly available at https://github.com/kerenfu/LFSOD-Survey .
Keren Fu, Yao Jiang 0002, Ge-Peng Ji, Tao Zhou 0002, Qijun Zhao, Deng-Ping Fan
Comput. Vis. Media5
2022 MEANet: Multi-modal edge-aware network for light field salient object detection
Yao Jiang 0002, Wenbo Zhang 0009, Keren Fu, Qijun Zhao
Neurocomputing4
2022 Discovering regression-detection bi-knowledge transfer for unsupervised cross-domain crowd counting
Yuting Liu 0004, Zheng Wang 0007, Miaojing Shi, Shin'ichi Satoh 0001, Qijun Zhao, Hongyu Yang 0002
Neurocomputing5
2022 Cascade wavelet transform based convolutional neural networks with application to image classification
Jieqi Sun, Qijun Zhao
Neurocomputing3
2022 SSPNet: Scale Selection Pyramid Network for Tiny Person Detection From UAV Images
abstract
With the increasing demand for search and rescue, it is highly demanded to detect objects of interest in large-scale images captured by unmanned aerial vehicles (UAVs), which is quite challenging due to extremely small scales of objects. Most existing methods employed a feature pyramid network (FPN) to enrich shallow layers’ features by combining deep layers’ contextual features. However, under the limitation of the inconsistency in gradient computation across different layers, the shallow layers in FPN are not fully exploited to detect tiny objects. In this article, we propose a scale selection pyramid network (SSPNet) for tiny person detection, which consists of three components: context attention module (CAM), scale enhancement module (SEM), and scale selection module (SSM). CAM takes account of context information to produce hierarchical attention heatmaps. SEM highlights features of specific scales at different layers, leading the detector to focus on objects of specific scales instead of vast backgrounds. SSM exploits adjacent layers’ relationships to fulfill suitable feature sharing between deep layers and shallow layers, thereby avoiding the inconsistency in gradient computation across different layers. Besides, we propose a weighted negative sampling (WNS) strategy to guide the detector to select more representative samples. Experiments on the TinyPerson benchmark show that our method outperforms other state-of-the-art (SOTA) detectors.
Mingbo Hong, Shuiwang Li, Feiyu Zhu 0001, Qijun Zhao
IEEE Geosci. Remote. Sens. Lett.5
2022 Siamese Network for RGB-D Salient Object Detection and Beyond
abstract
Existing RGB-D salient object detection (SOD) models usually treat RGB and depth as independent information and design separate networks for feature extraction from each. Such schemes can easily be constrained by a limited amount of training data or over-reliance on an elaborately designed training process. Inspired by the observation that RGB and depth modalities actually present certain commonality in distinguishing salient objects, a novel joint learning and densely cooperative fusion (JL-DCF) architecture is designed to learn from both RGB and depth inputs through a shared network backbone, known as the Siamese architecture. In this paper, we propose two effective components: joint learning (JL), and densely cooperative fusion (DCF). The JL module provides robust saliency feature learning by exploiting cross-modal commonality via a Siamese network, while the DCF module is introduced for complementary feature discovery. Comprehensive experiments using 5 popular metrics show that the designed framework yields a robust RGB-D saliency detector with good generalization. As a result, JL-DCF significantly advances the SOTAs by an average of ~2.0% (F-measure) across 7 challenging datasets. In addition, we show that JL-DCF is readily applicable to other related multi-modal detection tasks, including RGB-T SOD and video SOD, achieving comparable or better performance.
Keren Fu, Deng-Ping Fan, Ge-Peng Ji, Qijun Zhao, Jianbing Shen, Ce Zhu
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Learning residue-aware correlation filters and refining scale for real-time UAV tracking
Shuiwang Li, Yuting Liu 0004, Qijun Zhao, Ziliang Feng
Pattern Recognit.3
2022 Mutual-Guidance Transformer-Embedding Network for Video Salient Object Detection
abstract
Video salient object detection (VSOD) aims at locating the most attractive objects presented in video sequences by exploiting spatial and temporal cues. Previous methods mainly utilize convolutional neural networks (CNNs) to fuse or complement across RGB and optical flow cues via simple strategies. To take full advantage of CNNs and recently emerged Transformers, this letter proposes a novel mutual-guidance Transformer-embedding network, called MGT-Net, where a mutual-guidance multi-head attention mechanism (MGMA) explores more sophisticated long-range cross-modal interactions. Such a mechanism is designed into a new mutual-guidance Transformer (MGTrans) module that can propagate long-range contextual dependencies based on information of the other modality. To the best of our knowledge, MGT-Net is the first VSOD model that embeds Transformers as modules into CNNs for improved performance. Prior to MGTrans, we also propose and deploy a feature purification module (FPM) to purify noisy backbone features. Experimental results on five benchmark datasets demonstrate the state-of-the-art performance of MGT-Net.
Dingyao Min, Chao Zhang 0072, Yukang Lu, Keren Fu, Qijun Zhao
IEEE Signal Process. Lett.5
2021 Learning Residue-Aware Correlation Filters and Refining Scale Estimates with the GrabCut for Real-Time UAV Tracking
abstract
Unmanned aerial vehicle (UAV)-based tracking is attracting increasing attention and developing rapidly in applications such as agriculture, aviation, navigation, transportation and public security. Recently, discriminative correlation filters (DCF)-based trackers have stood out in UAV tracking community for their high efficiency and appealing robustness on a single CPU. However, due to limited onboard computation resources and other challenges the efficiency and accuracy of existing DCF-based approaches is still not satisfying. In this paper, inspired by residue representation, we exploit the residue nature inherent to videos and propose residue-aware correlation filters that show better convergence properties in filter learning. Moreover, we explore using segmentation by the GrabCut to improve the wildly adopted discriminative scale estimation in DCF-based trackers, which, as a mater of fact, greatly impacts the precision and accuracy of the trackers since accumulated scale error degrades the appearance model as online updating goes on. Extensive experiments are conducted on four UAV benchmarks, namely, UAV123@10fps, DTB70, UAVDT and Vistrone2018 (VisDrone2018-test-dev). The results show that our method achieves state-of-the-art performance.
Shuiwang Li, Yuting Liu 0004, Qijun Zhao, Ziliang Feng
3DV3
2021 RGB-D Salient Object Detection via 3D Convolutional Neural Networks
abstract
RGB-D salient object detection (SOD) recently has attracted increasing research interest and many deep learning methods based on encoder-decoder architectures have emerged. However, most existing RGB-D SOD models conduct feature fusion either in the single encoder or the decoder stage, which hardly guarantees sufficient cross-modal fusion ability. In this paper, we make the first attempt in addressing RGB-D SOD through 3D convolutional neural networks. The proposed model, named RD3D, aims at pre-fusion in the encoder stage and in-depth fusion in the decoder stage to effectively promote the full integration of RGB and depth streams. Specifically, RD3D first conducts pre-fusion across RGB and depth modalities through an inflated 3D encoder, and later provides in-depth feature fusion by designing a 3D decoder equipped with rich back-projection paths (RBPP) for leveraging the extensive aggregation ability of 3D convolutions. With such a progressive fusion strategy involving both the encoder and decoder, effective and thorough interaction between the two modalities can be exploited and boost the detection accuracy. Extensive experiments on six widely used benchmark datasets demonstrate that RD3D performs favorably against 14 state-of-the-art RGB-D SOD approaches in terms of four key evaluation metrics. Our code will be made publicly available: https://github.com/PPOLYpubki/RD3D.
Yi Zhang 0076, Keren Fu, Qijun Zhao, Hongwei Du 0004
AAAI5
2021 3D Face Point Cloud Super-Resolution Network
abstract
With the development of consumer-level depth sensors, 3D face point cloud data can be easily captured now. However, such data are often accompanied by low resolution, noise, and holes. At the same time, high-precision 3D scanners are bulky and can not be widely used in daily applications due to costs and inconvenience. To fill the gap between low and high resolution 3D faces, we propose a two-stage framework named the face point cloud super-resolution network (FPSRN) to recover high-resolution 3D face data from the low-resolution counterparts. As the human faces can be aligned into a unified coordinate system, we formulate point cloud super-resolution as a z-coordinate prediction problem. Cascaded auto-encoders are employed to retain both global structure and boundary information of different face regions during super-resolution. Compared with state- of-the-art point cloud completion methods and depth estimation methods, our method improves the Earth-Mover’s Distance (EMD) and the Root Mean Square Error (RMSE) metrics by 43% and 25%, respectively.
Feiyu Zhu 0001, Xiao Yang 0029, Qijun Zhao
IJCB4
2021 YakReID-103: A Benchmark for Yak Re-Identification
abstract
Precision livestock management requires animal traceability and disease trajectory, for which discriminating between or re-identifying individual animals is of significant importance. Existing re-identification (re-ID) methods are mostly proposed for persons and vehicles, compared with which animals are extraordinarily more challenging to be re-identified because of subtle visual differences between individuals. In this paper, we focus on image-based re-ID of yaks (Bos grunniens), which are indispensable livestock in local animal husbandry economy in Qinghai-Tibet Plateau. We establish the first yak re-ID dataset (called YakReID-103) which contains 2, 247 images of 103 different yaks with bounding box, direction-based pose, and identity annotations. Moreover, according to the characteristics of yaks, we modifiy several person re-ID and animal re-ID methods as baselines for yak re-ID. Experimental results of the baselines on YakReID-103 demonstrate the challenges in yak re-ID. We expect that the proposed benchmark will promote the research of animal biometrics and extend the application scope of re-ID techniques.
Qijun Zhao, Cuo Da, Liyuan Zhou, Suonan Jiancuo
IJCB2
2021 Learning Disentangled Representation for Fine-Grained Visual Categorization
Wenjie Dang, Shuiwang Li, Qijun Zhao
ICIG (1)3
2021 Equivalence of Correlation Filter and Convolution Filter in Visual Tracking
Shuiwang Li, Qijun Zhao, Ziliang Feng
ICIG (3)2
2021 Bird Keypoint Detection via Exploiting 2D Texture and 3D Geometric Features
Qijun Zhao, Pubu Danzeng
ICIG (3)2
2021 Watermark Faker: Towards Forgery of Digital Image Watermarking
abstract
Digital watermarking has been widely used to protect the copyright and integrity of multimedia data. Previous studies mainly focus on designing watermarking techniques that are robust to attacks of destroying the embedded watermarks. However, the emerging deep learning based image generation technology raises new open issues that whether it is possible to generate fake watermarked images for circumvention. In this paper, we make the first attempt to develop digital image watermark fakers by using generative adversarial learning. Suppose that a set of paired images of original and watermarked images generated by the targeted watermarker are available, we use them to train a watermark faker with U-Net as the backbone, whose input is an original image, and after a domain-specific preprocessing, it outputs a fake watermarked image. Our experiments show that the proposed watermark faker can effectively crack digital image watermarkers in both spatial and frequency domains, suggesting the risk of such forgery attacks.
Ruowei Wang, Chenguo Lin, Qijun Zhao, Feiyu Zhu 0001
ICME3
2021 BTS-Net: Bi-Directional Transfer-And-Selection Network for RGB-D Salient Object Detection
abstract
Depth information has been proved beneficial in RGB-D salient object detection (SOD). However, depth maps obtained often suffer from low quality and inaccuracy. Most existing RGB-D SOD models have no cross-modal interactions or only have unidirectional interactions from depth to RGB in their encoder stages, which may lead to inaccurate encoder features when facing low quality depth. To address this limitation, we propose to conduct progressive bidirectional interactions as early in the encoder stage, yielding a novel bi-directional transfer-and-selection network named BTS-Net, which adopts a set of bi-directional transfer-and-selection (BTS) modules to purify features during encoding. Based on the resulting robust encoder features, we also design an effective light-weight group decoder to achieve accurate final saliency prediction. Comprehensive experiments on six widely used datasets demonstrate that BTS-Net surpasses 16 latest state-of-the-art approaches in terms of four key metrics.
Wenbo Zhang 0009, Yao Jiang 0002, Keren Fu, Qijun Zhao
ICME4
2021 Towards Silhouette-Aware Human Detection in Depth Images
abstract
Detecting humans in depth images attracts increasing attention thanks to the advantage of depth modality in privacy protection. However, this task is challenging because the number of available training data is limited to date and the depth images, unlike RGB images, are short of rich texture features. Although many image synthesis methods and deep learning methods have been proposed and proven successful, especially for RGB images, it is unsatisfactory to directly apply them to depth images because of the intrinsical differences between the modalities. In view of that silhouette of humans becomes an essential discriminative cue in depth images in the absence of texture information, which is not well utilized by existing methods, in this paper, we thus propose a silhouette-aware network (SAN) to train the detection model and a depth image synthesis method that represses spurious silhouette to augment the training data. Besides, to further increase the diversity of training data, we collect a dataset of scene depth images (SDI), including both indoor and outdoor scenes, as background images when synthesizing training data. Experimental results show that (i) our proposed synthesis method can generate more realistic depth images and thus benefits the training of detection models, (ii) our collected SDI dataset can effectively enhance data diversity and thus improves the effectiveness of the obtained detection models, and (iii) our proposed silhouette-aware network (SAN) can effectively boost the human detection accuracy. Our dataset is available at https://pan.baidu.com/s/13hpuziavBNjS8KATClpXww, password: r9id
Shuiwang Li, Qijun Zhao
IJCNN3
2021 Weakly Labeled Semi-Supervised Sound Event Detection with Multi-Scale Residual Attention
abstract
Different sound events have different time-frequency scale characteristics, which are useful for sound event detection (SED), but not yet effectively exploited. In this paper, we aim to adaptively select multi-scale feature information that is conducive to classification of sound events. We propose a novel module, namely multi-scale residual attention (MSRA), which is composed of multi-scale residual convolutional block and selective multiscale attention block. Multi-scale residual convolution block extracts features at multiple scales, among which selective multiscale attention block adaptively selects the features that are helpful for event classification. Experimental results prove that our method outperforms the state-of-the-art model by 3.7% on Task 4 of the DCASE 2018 Challenge dataset.
Maolin Tang, Qijun Zhao, Zhengxi Liu
IJCNN2
2021 Depth Quality-Inspired Feature Manipulation for Efficient RGB-D Salient Object Detection
abstract
RGB-D salient object detection (SOD) recently has attracted increasing research interest by benefiting conventional RGB SOD with extra depth information. However, existing RGB-D SOD models often fail to perform well in terms of both efficiency and accuracy, which hinders their potential applications on mobile devices and real-world problems. An underlying challenge is that the model accuracy usually degrades when the model is simplified to have few parameters. To tackle this dilemma and also inspired by the fact that depth quality is a key factor influencing the accuracy, we propose a novel depth quality-inspired feature manipulation (DQFM) process, which is efficient itself and can serve as a gating mechanism for filtering depth features to greatly boost the accuracy. DQFM resorts to the alignment of low-level RGB and depth features, as well as holistic attention of the depth stream to explicitly control and enhance cross-modal fusion. We embed DQFM to obtain an efficient light-weight model called DFM-Net, where we also design a tailored depth backbone and a two-stage decoder for further efficiency consideration. Extensive experimental results demonstrate that our DFM-Net achieves state-of-the-art accuracy when comparing to existing non-efficient models, and meanwhile runs at 140ms on CPU (2.2x faster than the prior fastest efficient model) with only ~8.5Mb model size (14.9% of the prior lightest). Our code will be available at https://github.com/zwbx/DFM-Net.
Wenbo Zhang 0009, Ge-Peng Ji, Zhuo Wang 0004, Keren Fu, Qijun Zhao
ACM Multimedia5
2020 Weakly-Supervised Reconstruction of 3D Objects with Large Shape Variation from Single In-the-Wild Images
Shichen Sun, Zhengbang Zhu, Xiaowei Dai, Qijun Zhao, Jing Li 0060
ACCV (1)4
2020 JL-DCF: Joint Learning and Densely-Cooperative Fusion Framework for RGB-D Salient Object Detection
abstract
This paper proposes a novel joint learning and densely-cooperative fusion (JL-DCF) architecture for RGB-D salient object detection. Existing models usually treat RGB and depth as independent information and design separate networks for feature extraction from each. Such schemes can easily be constrained by a limited amount of training data or over-reliance on an elaborately-designed training process. In contrast, our JL-DCF learns from both RGB and depth inputs through a Siamese network. To this end, we propose two effective components: joint learning (JL), and densely-cooperative fusion (DCF). The JL module provides robust saliency feature learning, while the latter is introduced for complementary feature discovery. Comprehensive experiments on four popular metrics show that the designed framework yields a robust RGB-D saliency detector with good generalization. As a result, JL-DCF significantly advances the top-1 D3Net model by an average of ~1.9% (S-measure) across six challenging datasets, showing that the proposed framework offers a potential solution for real-world applications and could provide more insight into the cross-modality complementarity task. The code will be available at https://github.com/kerenfu/JLDCF/.
Keren Fu, Deng-Ping Fan, Ge-Peng Ji, Qijun Zhao
CVPR4
2020 IF-GAN: Generative Adversarial Network for Identity Preserving Facial Image Inpainting and Frontalization
abstract
Face recognition has seen rapid development and widespread usage in recent years. However, recognizing faces with occlusions or in large pose variations still poses great challenges for existing face recognition algorithms. While these two obstructions frequently occur in real-world applications simultaneously, many recent research efforts can only solve one of these two problems. To address this issue, in this paper, we propose an identity preserving two-stage generative adversarial network that can simultaneously complete the task of de-occlusion and face frontalization. Unlike previous works that can only inpaint synthetic rectangular occlusion which is unlikely to occur in real-life scenarios, we employ partial convolutions in our face inpainting stage to handle the realistic irregular occlusions. For face frontalization, we use a dualpathway structure that processes global facial shape and local facial fiducial region details separately, allowing the network to gain enough semantic information to faithfully reconstruct the frontal facial image. Both stages are supervised on both pixel and feature levels such that the network can produce photo-realistic yet identity preserving unoccluded frontal facial images. Qualitative and quantitative experiments demonstrate that the proposed approach is able to improve the performance of existing face recognition systems by restoring identity information contaminated by occlusions and pose variations.
Kunjian Li, Qijun Zhao
FG2
2020 Detection Features as Attention (Defat): A Keypoint-Free Approach to Amur Tiger Re-Identification
abstract
Automatically identifying animals in camera-trap images has attracted increasing attention due to its valuable potential in wildlife conservation. A typical pipeline of existing methods includes separated animal detection and re-identification modules, and state-of-the-art re-identification methods either use annotated keypoints of animals to extract robust features, or employ extra branches to learn multiple features. In this paper, in contrast, we propose a keypoint-free approach to Amur tiger re-identification by exploiting the feature maps extracted by the detection module to help the re-identification module learn more effective features. We devise a detection-features-as-attention (DeFAt) module, which generates an additive mask for the input image based on the detection feature maps. We experimentally show that using the masked image the re-identification module lays more attention on the tiger region in the image, while the distraction by the messy background is removed to some extent. Our evaluation results prove that the proposed DeFAt module can effectively improve the Amur tiger re-identification accuracy when key-point annotations are not available.
Xinhua Cheng, Jianing Zhu, Qijun Zhao
ICIP5
2020 Towards Unsupervised Crowd Counting via Regression-Detection Bi-knowledge Transfer
abstract
Unsupervised crowd counting is a challenging yet not largely explored task. In this paper, we explore it in a transfer learning setting where we learn to detect and count persons in an unlabeled target set by transferring bi-knowledge learnt from regression- and detection-based models in a labeled source set. The dual source knowledge of the two models is heterogeneous and complementary as they capture different modalities of the crowd distribution. We formulate the mutual transformations between the outputs of regression- and detection-based models as two scene-agnostic transformers which enable knowledge distillation between the two models. Given the regression- and detection-based models and their mutual transformers learnt in the source, we introduce an iterative self-supervised learning scheme with regression-detection bi-knowledge transfer in the target. Extensive experiments on standard crowd counting benchmarks, ShanghaiTech, UCF_CC_50, and UCF_QNRF demonstrate a substantial improvement of our method over other state-of-the-arts in the transfer learning setting.
Yuting Liu 0004, Zheng Wang 0007, Miaojing Shi, Shin'ichi Satoh 0001, Qijun Zhao, Hongyu Yang 0002
ACM Multimedia5
2020 3D face reconstruction from mugshots: Application to arbitrary view face recognition
Huan Tu, Feng Liu 0037, Qijun Zhao, Anil K. Jain 0001
Neurocomputing4
2020 Robust and high-security fingerprint recognition system using optical coherence tomography
Feng Liu 0013, Guojie Liu, Qijun Zhao, LinLin Shen
Neurocomputing3
2020 Asymmetric discriminative correlation filters for visual tracking
abstract
Discriminative correlation filters (DCF) are efficient in visual tracking and have advanced the field significantly. However, the symmetry of correlation (or convolution) operator results in computational problems and does harm to the generalized translation equivariance. The former problem has been approached in many ways, whereas the latter one has not been well recognized. In this paper, we analyze the problems with the symmetry of circular convolution and propose an asymmetric one, which as a generalization of the former has a weak generalized translation equivariance property. With this operator, we propose a tracker called the asymmetric discriminative correlation filter (ADCF), which is more sensitive to translations of targets. Its asymmetry allows the filter and the samples to have different sizes. This flexibility makes the computational complexity of ADCF more controllable in the sense that the number of filter parameters will not grow with the sample size. Moreover, the normal matrix of ADCF is a block matrix with each block being a two-level block Toeplitz matrix. With this well-structured normal matrix, we design an algorithm for multiplying an N × N two-level block Toeplitz matrix by a vector with time complexity O ( N log N ) and space complexity O ( N ), instead of O ( N 2 ). Unlike DCF-based trackers, introducing spatial or temporal regularization does not increase the essential computational complexity of ADCF. Comparative experiments are performed on a synthetic dataset and four benchmarks, including OTB-2013, OTB-2015, VOT-2016, and Temple-Color, and the results show that our method achieves state-of-the-art visual tracking performance.
Shuiwang Li, Qianbo Jiang, Qijun Zhao, Ziliang Feng
Frontiers Inf. Technol. Electron. Eng.3
2020 Joint Face Alignment and 3D Face Reconstruction with Application to Face Recognition
abstract
Face alignment and 3D face reconstruction are traditionally accomplished as separated tasks. By exploring the strong correlation between 2D landmarks and 3D shapes, in contrast, we propose a joint face alignment and 3D face reconstruction method to simultaneously solve these two problems for 2D face images of arbitrary poses and expressions. This method, based on a summation model of 3D faces and cascaded regression in 2D and 3D shape spaces, iteratively and alternately applies two cascaded regressors, one for updating 2D landmarks and the other for 3D shape. The 3D shape and the landmarks are correlated via a 3D-to-2D mapping matrix, which is updated in each iteration to refine the location and visibility of 2D landmarks. Unlike existing methods, the proposed method can fully automatically generate both pose-and-expression-normalized (PEN) and expressive 3D faces and localize both visible and invisible 2D landmarks. Based on the PEN 3D faces, we devise a method to enhance face recognition accuracy across poses and expressions. Both linear and nonlinear implementations of the proposed method are presented and evaluated in this paper. Extensive experiments show that the proposed method can achieve the state-of-the-art accuracy in both face alignment and 3D face reconstruction, and benefit face recognition owing to its reconstructed PEN 3D face.
Feng Liu 0037, Qijun Zhao, Xiaoming Liu 0002, Dan Zeng 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Point in, Box Out: Beyond Counting Persons in Crowds
abstract
Modern crowd counting methods usually employ deep neural networks (DNN) to estimate crowd counts via density regression. Despite their significant improvements, the regression-based methods are incapable of providing the detection of individuals in crowds. The detection-based methods, on the other hand, have not been largely explored in recent trends of crowd counting due to the needs for expensive bounding box annotations. In this work, we instead propose a new deep detection network with only point supervision required. It can simultaneously detect the size and location of human heads and count them in crowds. We first mine useful person size information from point-level annotations and initialize the pseudo ground truth bounding boxes. An online updating scheme is introduced to refine the pseudo ground truth during training; while a locally-constrained regression loss is designed to provide additional constraints on the size of the predicted boxes in a local neighborhood. In the end, we propose a curriculum learning strategy to train the network from images of relatively accurate and easy pseudo ground truth first. Extensive experiments are conducted in both detection and counting tasks on several standard benchmarks, e.g. ShanghaiTech, UCF_CC_50, WiderFace, and TRANCOS datasets, and the results show the superiority of our method over the state-of-the-art.
Yuting Liu 0004, Miaojing Shi, Qijun Zhao
CVPR3
2019 Learning to Focus and Discriminate for Fine-Grained Classification
abstract
Existing state-of-the-art fine-grained classification methods usually use separated networks for discriminative region localization and feature learning/classification, and are thus complicated to implement and optimize. In this paper, we aim to provide a compact solution by deepening the collaboration between the region localization, feature learning and classification modules during the learning process of fine-grained classification. We thus propose a method that can learn to simultaneously localize discriminative regions and extract discriminative features by exploring the localization ability of classification convolutional neural networks and joint optimization of different modules. Our method, while being built upon a single backbone network and trained with only softmax losses, achieves state-of-the-art performance on three benchmark fine-grained datasets, which proves that our method is simple but effective for fine-grained classification.
Zhicong Feng, Keren Fu, Qijun Zhao
ICIP3
2019 Pooling Scores of Neighboring Points for Improved 3D Point Cloud Segmentation
abstract
3D point cloud segmentation has been rapidly advanced by exploring features of neighboring points. However, existing methods still suffer from ambiguous features, especially for the points in junction regions. In this paper, we show that such problem can be alleviated by utilizing the segmentation scores of neighboring points. We thus propose an attention-based score refinement module, which can be easily integrated with existing 3D point cloud segmentation networks and improve the segmentation accuracy. We demonstrate the effectiveness of our proposed method with extensive experiments on two challenging datasets (i.e., ShapeNet and ScanNet).
Weihao Zhou, Qijun Zhao
ICIP4
2019 Dense Correspondence of 3d Facial Point Clouds Via Neural Network Fitting
abstract
3D face dense correspondence is an important and challenging problem in 3D face analysis. Previous methods usually solve this problem by applying transformations to align different 3D faces and are thus constrained by the employed specific transformations. In this paper, instead, we approach the problem as a surface fitting and re-sampling problem. Modelling the point cloud of a 3D face as a surface, we use a neural network to fit the surface, and evaluate the obtained network at pre-specified sampling points that are defined based on prior knowledge of face structure. We evaluate the proposed method on the BU-3DFE database and prove its effectiveness both qualitatively and quantitatively.
Weihao Zhou, Qijun Zhao
ICIP4
2019 Distinguishing Individual Red Pandas from Their Faces
Qijun Zhao, Zhihe Zhang, Rong Hou
PRCV (2)2
2019 Combined training strategy for low-resolution face recognition with limited application-specific data
abstract
Application‐specific data for certain biometric applications are often not sufficiently available. The authors present a solution for face recognition with limited application‐specific data. Existing methods often use a classifier with convolutional neural networks (CNNs) as feature extractors. The CNNs are trained with massive general (i.e. not application specific) data and the classifier is trained with application‐specific data. Alternatively, the authors propose a combined training strategy to train the classifier on a balanced mixture of general and application‐specific data, such that the recognition performance is maximised. The proposed method largely alleviates the needs for application‐specific data. To prove its effectiveness, they apply the proposed method to low‐resolution face recognition. Specifically, they use the heterogeneous joint Bayesian (HJB) classifier that is capable of comparing features from the same modality but with different characteristics. To further boost performance, the authors augment the training data by pre‐processing it to resemble application‐specific data. They conducted extensive experiments on challenging datasets, namely, SCface and COX. The results show that the proposed method improves the true match rate on SCface at a false match rate of 10% by ∼11% and the true match rate on COX at a false match rate of 1% by ∼12%.
Dan Zeng 0002, Luuk J. Spreeuwers, Raymond N. J. Veldhuis, Qijun Zhao
IET Image Process.4
2019 Deepside: A general deep framework for salient object detection
Keren Fu, Qijun Zhao, Irene Y. H. Gu, Jie Yang 0002
Neurocomputing2
2019 A variational image segmentation method exploring both intensity means and texture patterns
Qijun Zhao, Xiangchu Feng, Weiwei Wang 0005, Renrui Zhang, An Yan 0004
Signal Process. Image Commun.2
2019 Refinet: A Deep Segmentation Assisted Refinement Network for Salient Object Detection
abstract
Compared to conventional saliency detection by handcrafted features, deep convolutional neural networks (CNNs) recently have been successfully applied to saliency detection field with superior performance on locating salient objects. However, due to repeated sub-sampling operations inside CNNs such as pooling and convolution, many CNN-based saliency models fail to maintain fine-grained spatial details and boundary structures of objects. To remedy this issue, this paper proposes a novel end-to-end deep learning-based refinement model named Refinet, which is based on fully convolutional network augmented with segmentation hypotheses. Intermediate saliency maps that are edge-aware are computed from segmentation-based pooling and then feed to a two-tier fully convolutional network for effective fusion and refinement, leading to more precise object details and boundaries. In addition, the resolution of feature maps in the proposed Refinet is carefully designed to guarantee sufficient boundary clarity of the refined saliency output. Compared to widely employed dense conditional random field, Refinet is able to enhance coarse saliency maps generated by existing models with more accurate spatial details, and its effectiveness is demonstrated by experimental results on seven benchmark datasets.
Keren Fu, Qijun Zhao, Irene Y. H. Gu
IEEE Trans. Multim.2
2018 Disentangling Features in 3D Face Shapes for Joint Face Reconstruction and Recognition
abstract
This paper proposes an encoder-decoder network to disentangle shape features during 3D face reconstruction from single 2D images, such that the tasks of reconstructing accurate 3D face shapes and learning discriminative shape features for face recognition can be accomplished simultaneously. Unlike existing 3D face reconstruction methods, our proposed method directly regresses dense 3D face shapes from single 2D images, and tackles identity and residual (i.e., non-identity) components in 3D face shapes explicitly and separately based on a composite 3D face shape model with latent representations. We devise a training process for the proposed network with a joint loss measuring both face identification error and 3D face shape reconstruction error. To construct training data we develop a method for fitting 3D morphable model (3DMM) to multiple 2D images of a subject. Comprehensive experiments have been done on MICC, BU3DFE, LFW and YTF databases. The results show that our method expands the capacity of 3DMM for capturing discriminative shape features and facial detail, and thus outperforms existing methods both in 3D face reconstruction accuracy and in face recognition accuracy.
Feng Liu 0013, Ronghang Zhu, Dan Zeng 0002, Qijun Zhao, Xiaoming Liu 0002
CVPR4
2018 Evaluation of Dense 3D Reconstruction from 2D Face Images in the Wild
abstract
This paper investigates the evaluation of dense 3D face reconstruction from a single 2D image in the wild. To this end, we organise a competition that provides a new benchmark dataset that contains 2000 2D facial images of 135 subjects as well as their 3D ground truth face scans. In contrast to previous competitions or challenges, the aim of this new benchmark dataset is to evaluate the accuracy of a 3D dense face reconstruction algorithm using real, accurate and high-resolution 3D ground truth face scans. In addition to the dataset, we provide a standard protocol as well as a Python script for the evaluation. Last, we report the results obtained by three state-of-the-art 3D face reconstruction systems on the new benchmark dataset. The competition is organised along with the 2018 13th IEEE Conference on Automatic Face & Gesture Recognition.
Zhenhua Feng 0001, Patrik Huber 0001, Josef Kittler, Peter J. B. Hancock, Xiaojun Wu 0001, Qijun Zhao, Willem P. Koppen, Matthias Rätsch
FG6
2018 Landmark-Based 3D Face Reconstruction from an Arbitrary Number of Unconstrained Images
abstract
In this paper, we propose a novel method for reconstructing 3D faces from 2D images. The method is characterized in three aspects. (i) It utilizes only geometric cues in the input images, i.e., 2D facial landmarks. (ii) It works for an arbitrary number of unconstrained images, both single and multiple images. (iii) It can effectively exploit complementary information in multiple images of varying poses and expressions. The method is implemented based on cascaded regression in shape space. We have evaluated the method on three databases and observed from the experimental results that (i) the reconstruction error is reduced as more images of different poses are used, (ii) the proposed method can obtain comparable reconstruction results by using state-of-the-art automated methods to detect the 2D landmarks, and (iii) the proposed method is robust to variations in facial expressions and image qualities.
Wan Tian, Feng Liu 0037, Qijun Zhao
FG3
2018 On Mugshot-based Arbitrary View Face Recognition
abstract
Despite the wide usage of mugshot images in forensic applications, they are underutilized in existing automated face recognition systems. In this paper, we propose a novel mugshot-based arbitrary view face recognition method. Our approach reconstructs full 3D faces via cascaded regression in shape space with efficient seamless texture recovery. Unlike existing methods, it makes full use of the frontal and profile views available in mugshot images, and thus generates accurate and realistic 3D faces. Multi-view face images are synthesized from the reconstructed 3D faces to enlarge the gallery so that arbitrary view faces can be better recognized. Evaluation experiments were conducted on BFM and Multi-PIE databases by using state-of-the-art deep learning (DL) based face matchers. The results demonstrate the effectiveness of our proposed method and show that DL-based face matchers can benefit from mugshot images and the reconstructed 3D faces, especially for recognizing large off-angle faces.
Feng Liu 0037, Huan Tu, Qijun Zhao, Anil K. Jain 0001
ICPR4
2017 Multi-dim: A multi-dimensional face database towards the application of 3D technology in real-world scenarios
abstract
Three-dimensional (3D) faces are increasingly utilized in many face-related tasks. Despite the promising improvement achieved by 3D face technology, it is still hard to thoroughly evaluate the performance and effect of 3D face technology in real-world applications where variations frequently occur in pose, illumination, expression and many other factors. This is due to the lack of benchmark databases that contain both high precision full-view 3D faces and their 2D face images/videos under different conditions. In this paper, we present such a multi-dimensional face database (namely Multi-Dim) of high precision 3D face scans, high definition photos, 2D still face images with varying pose and expression, low quality 2D surveillance video clips, along with ground truth annotations for them. Based on this Multi-Dim face database, extensive evaluation experiments have been done with state-of-the-art baseline methods for constructing 3D morphable model, reconstructing 3D faces from single images, 3D-assisted pose normalization for face verification, and 3D-rendered multiview gallery for face identification. Our results show that 3D face technology does help in improving unconstrained 2D face recognition when the probe 2D face images are of reasonable quality, whereas it deteriorates rather than improves the face recognition accuracy when the probe 2D face images are of poor quality. We will make Multi-Dim freely available to the community for the purpose of advancing the 3D-based unconstrained 2D face recognition and related techniques towards real-world applications.
Feng Liu 0013, Qijun Zhao
IJCB5
2017 2.5D cascaded regression for robust facial landmark detection
abstract
In this paper, we propose a 2.5D Cascaded Regression approach for accurately and robustly locating facial landmarks on RGB-D data. Instead of detecting facial landmarks on texture and depth images separately, the proposed method alternately applies depth-based and texture-based regressors to compute the necessary increments to the estimated landmarks so that they are gradually moved towards their true positions. This way, depth information is better explored through close interaction with texture information, and they together improve the facial landmark detection accuracy. Moreover, thanks to the robustness of depth information to illumination variations and its capacity of capturing the deformations caused by pose and expression changes, the proposed method has good robustness to pose, illumination and expression (PIE) variations. We have extensively evaluated the effectiveness of depth information and compared the proposed method with state-of-the-art texture-based and RGB-D-based methods on three publicly accessible databases, i.e., LIDF, EURECOM and Curtin-Faces. The evaluation results validate the superiority of our approach in utilizing depth information for accurately detecting facial landmarks under challenging conditions with obvious PIE variations.
Jinwen Xu, Qijun Zhao
IJCB2
2017 Encrypted domain matching of fingerprint minutia cylinder-code (MCC) with l1 minimization
Eryun Liu, Qijun Zhao
Neurocomputing2
2017 Examplar coherent 3D face reconstruction from forensic mugshot database
Dan Zeng 0002, Qijun Zhao, Shuqin Long, Jing Li 0060
Image Vis. Comput.2
2017 On 3D face reconstruction via cascaded regression in shape space
abstract
Cascaded regression has been recently applied to reconstruct 3D faces from single 2D images directly in shape space, and has achieved state-of-the-art performance. We investigate thoroughly such cascaded regression based 3D face reconstruction approaches from four perspectives that are not well been studied: (1) the impact of the number of 2D landmarks; (2) the impact of the number of 3D vertices; (3) the way of using standalone automated landmark detection methods; (4) the convergence property. To answer these questions, a simplified cascaded regression based 3D face reconstruction method is devised. This can be integrated with standalone automated landmark detection methods and reconstruct 3D face shapes that have the same pose and expression as the input face images, rather than normalized pose and expression. An effective training method is also proposed by disturbing the automatically detected landmarks. Comprehensive evaluation experiments have been carried out to compare to other 3D face reconstruction methods. The results not only deepen the understanding of cascaded regression based 3D face reconstruction approaches, but also prove the effectiveness of the proposed method.
Feng Liu 0013, Dan Zeng 0002, Jing Li 0060, Qijun Zhao
Frontiers Inf. Technol. Electron. Eng.4
2016 Joint Face Alignment and 3D Face Reconstruction
Feng Liu 0013, Dan Zeng 0002, Qijun Zhao, Xiaoming Liu 0002
ECCV (5)3
2016 Unseen head pose prediction using dense multivariate label distribution
abstract
Accurate head poses are useful for many face-related tasks such as face recognition, gaze estimation, and emotion analysis. Most existing methods estimate head poses that are included in the training data (i.e., previously seen head poses). To predict head poses that are not seen in the training data, some regression-based methods have been proposed. However, they focus on estimating continuous head pose angles, and thus do not systematically evaluate the performance on predicting unseen head poses. In this paper, we use a dense multivariate label distribution (MLD) to represent the pose angle of a face image. By incorporating both seen and unseen pose angles into MLD, the head pose predictor can estimate unseen head poses with an accuracy comparable to that of estimating seen head poses. On the Pointing’04 database, the mean absolute errors of results for yaw and pitch are 4.01° and 2.13°, respectively. In addition, experiments on the CAS-PEAL and CMU Multi-PIE databases show that the proposed dense MLD-based head pose estimation method can obtain the state-of-the-art performance when compared to some existing methods.
Gaoli Sang, Hu Chen 0002, Ge Huang, Qijun Zhao
Frontiers Inf. Technol. Electron. Eng.4
2014 Detecting soft shadows in a single outdoor image: From local edge-based models to global constraints
Yanli Liu 0002, Qijun Zhao, Hongyu Yang 0002
Comput. Graph.3
2014 Fingerprint orientation field reconstruction by weighted discrete cosine transform
Manhua Liu, Qijun Zhao
Inf. Sci.3
2013 Feature extraction using two-dimensional neighborhood margin and variation embedding
Quanxue Gao, Xiujuan Hao, Qijun Zhao, Weiguo Shen, Jingjie Ma
Comput. Vis. Image Underst.3
2012 Model Based Separation of Overlapping Latent Fingerprints
abstract
Latent fingerprints lifted from crime scenes often contain overlapping prints, which are difficult to separate and match by state-of-the-art fingerprint matchers. A few methods have been proposed to separate overlapping fingerprints to enable fingerprint matchers to successfully match the component fingerprints. These methods are limited by the accuracy of the estimated orientation field, which is not reliable for poor quality overlapping latent fingerprints. In this paper, we improve the robustness of overlapping fingerprints separation, particularly for low quality images. Our algorithm reconstructs the orientation fields of component prints by modeling fingerprint orientation fields. In order to facilitate this, we utilize the orientation cues of component fingerprints, which are manually marked by fingerprint examiners. This additional markup is acceptable in forensics, where the first priority is to improve the latent matching accuracy. The effectiveness of the proposed method has been evaluated not only on simulated overlapping prints, but also on real overlapped latent fingerprint images. Compared with available methods, the proposed algorithm is more effective in separating poor quality overlapping fingerprints and enhancing the matching accuracy of overlapping fingerprints.
Qijun Zhao, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.1
2011 3D to 2D fingerprints: Unrolling and distortion correction
abstract
Touchless 3D fingerprint sensors can capture both 3D depth information and albedo images of the finger surface. Compared with 2D fingerprint images acquired by traditional contact-based fingerprint sensors, the 3D fingerprints are generally free from the distortion caused by non-uniform pressure and undesirable motion of the finger. Several unrolling algorithms have been proposed for virtual rolling of 3D fingerprints to obtain 2D equivalent fingerprints, so that they can be matched with the legacy 2D fingerprint databases. However, available unrolling algorithms do not consider the impact of distortion that is typically present in the legacy 2D fingerprint images. In this paper, we conduct a comparative study of representative unrolling algorithms and propose an effective approach to incorporate distortion into the unrolling process. The 3D fingerprint database was acquired by using a 3D fingerprint sensor being developed by the General Electric Global Research. By matching the 2D equivalent fingerprints with the corresponding 2D fingerprints collected with a commercial contact-based fingerprint sensor, we show that the compatibility between the 2D unrolled fingerprints and the traditional contact-based 2D fingerprints is improved after incorporating the distortion into the unrolling process.
Qijun Zhao, Anil K. Jain 0001, Gil Abramovich
IJCB1
2011 Fast kernel Fisher discriminant analysis via approximating the kernel principal component analysis
Qin Li 0001, Jane You, Qijun Zhao
Neurocomputing4
2011 A novel hierarchical fingerprint matching approach
Feng Liu 0013, Qijun Zhao, David Zhang 0001
Pattern Recognit.2
2011 Facial expression recognition on multiple manifolds
Qijun Zhao, David Zhang 0001
Pattern Recognit.2
2011 Quantitative analysis of human facial beauty using geometric features
David Zhang 0001, Qijun Zhao, Fangmei Chen
Pattern Recognit.2
2010 A comparative study on quality assessment of high resolution fingerprint images
abstract
High resolution fingerprint images have been increasingly used in fingerprint recognition. They can provide more fine features (e.g. pores) than standard fingerprint images to improve the recognition accuracy. It is however still an open issue whether or not existing quality assessment methods are suitable for high resolution fingerprint images. This paper compares some typical quality indexes by analyzing the correlation between them and their prediction ability on minutia-based and pore-based high resolution fingerprint recognition accuracy. Experimental results show that the indexes based on ridge orientation are more effective for high resolution fingerprint recognition systems.
Qijun Zhao, Feng Liu 0013, Lei Zhang 0006, David Zhang 0001
ICIP1
2010 Fingerprint Pore Matching Based on Sparse Representation
abstract
This paper proposes an improved direct fingerprint pore matching method. It measures the differences between pores by using the sparse representation technique. The coarse pore correspondences are then established and weighted based on the obtained differences. The false correspondences among them are finally removed by using the weighted RANSAC algorithm. Experimental results have shown that the proposed method can greatly improve the accuracy of existing methods.
Feng Liu 0013, Qijun Zhao, Lei Zhang 0006, David Zhang 0001
ICPR2
2010 Data Classification on Multiple Manifolds
abstract
Unlike most previous manifold-based data classification algorithms assume that all the data points are on a single manifold, we expect that data from different classes may reside on different manifolds of possible different dimensions. Therefore, better classification accuracy would be achieved by modeling the data by multiple manifolds each corresponding to a class. To this end, a general framework for data classification on multiple manifolds is presented. The manifolds are firstly learned for each class separately, and a stochastic optimization algorithm is then employed to get the near optimal dimensionality of each manifold from the classification viewpoint. Then, classification is performed under a newly defined minimum reconstruction error based classifier. Our method could be easily extended by involving various manifold learning methods and searching strategies. Experiments on both synthetic data and databases of facial expression images show the effectiveness of the proposed multiple manifold based approach.
Qijun Zhao, David Zhang 0001
ICPR2
2010 Parallel versus Hierarchical Fusion of Extended Fingerprint Features
abstract
Extended fingerprint features such as pores, dots and incipient ridges have been increasingly attracting attention from researchers and engineers working on automatic fingerprint recognition systems. A variety of methods have been proposed to combine these features with the traditional minutiae features. This paper comparatively analyses the parallel and hierarchical fusion approaches on a high resolution fingerprint image dataset. Based on the results, a novel and more effective hierarchical approach is presented for combining minutiae, pores, dots and incipient ridges.
Qijun Zhao, Feng Liu 0013, Lei Zhang 0006, David Zhang 0001
ICPR1
2010 High resolution partial fingerprint alignment using pore-valley descriptors
Qijun Zhao, David Zhang 0001, Lei Zhang 0006, Nan Luo
Pattern Recognit.1
2010 Adaptive fingerprint pore modeling and extraction
Qijun Zhao, David Zhang 0001, Lei Zhang 0006, Nan Luo
Pattern Recognit.1
2009 Curvature and singularity driven diffusion for oriented pattern enhancement with singular points
abstract
Oriented patterns, e.g. fingerprints, consist of smoothly varying flow-like patterns, together with important singular points (i.e. cores and deltas) where the orientation changes abruptly. Gabor filters and anisotropic diffusion methods have been widely used to enhance oriented patterns. However, none of them can well cope with regions of varying curvatures or regions surrounding singular points. By incorporating the ridge curvatures and the singularities into the diffusion model, we propose a new diffusion method to better exploit the global characteristics of oriented patterns. Specifically, we first locate the singular points, and regularize the estimated orientation field by using a singularity driven nonlinear diffusion process. We then enhance the oriented patterns by applying an oriented diffusion process which is driven by the curvature and singularity. Experiments on synthetic data and real fingerprint images validated that the proposed method is capable of consistently enhancing oriented patterns while well preserving the ridge structures in singular regions.
Qijun Zhao, Lei Zhang 0006, David Zhang 0001, Wenyi Huang
CVPR1
2008 Adaptive pore model for fingerprint pore extraction
abstract
Sweat pores have been recently employed for automated fingerprint recognition, in which the pores are usually extracted by using a computationally expensive skeletonization method or a unitary scale isotropic pore model. In this paper, however, we show that real pores are not always isotropic. To accurately and robustly extract pores, we propose an adaptive anisotropic pore model, whose parameters are adjusted adaptively according to the fingerprint ridge direction and period. The fingerprint image is partitioned into blocks and a local pore model is determined for each block. With the local pore model, a matched filter is used to extract the pores within each block. Experiments on a high resolution (1200dpi) fingerprint dataset are performed and the results demonstrate that the proposed pore model and pore extraction method can locate pores more accurately and robustly in comparison with other state-of-the-art pore extractors.
Qijun Zhao, Lei Zhang 0006, David Zhang 0001, Nan Luo, Jing Bao
ICPR1
2007 PCA-based web page watermarking
Qijun Zhao
Pattern Recognit.1
2006 A Direct Evolutionary Feature Extraction Algorithm for Classifying High Dimensional Data
Qijun Zhao, David Zhang 0001, Hongtao Lu 0001
AAAI1
2006 Parsimonious Feature Extraction Based on Genetic Algorithms and Support Vector Machines
Qijun Zhao, Hongtao Lu 0001, David Zhang 0001
ISNN (1)1
2006 A fast evolutionary pursuit algorithm based on linearly combining vectors
Qijun Zhao, Hongtao Lu 0001, David Zhang 0001
Pattern Recognit.1
2005 A PCA-based watermarking scheme for tamper-proof of web pages
Qijun Zhao
Pattern Recognit.1