Qiangqiang Zhou

dblp:161/2504 · DBLP profile ↗
← Back
30ranked-venue papers
1as first author
22since 2021 · last 2026
0000-0002-5717-3290ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 14 · 11 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 An industrial informatics-oriented multi-scale convolutional Mamba with multi-frequency attention for robust medical image segmentation
Yugen Yi, Wei Zhou 0003, Qiangqiang Zhou, Aiwen Jiang, Naixue Xiong, Yingkui Du, Xiaomei Huang
Eng. Appl. Artif. Intell.4
2025 BIDP: Brain-Inspired Dual-Process CNN-Transformer for Salient Object Detection
Wenyi Wu, Chen Liao, Qiangqiang Zhou, Dandan Zhu 0001, Xinping Rao
CGI (2)3
2025 A Camouflaged Object Detection Network with Global Cross-Space Perception and Flexible Local Feature Refinement
Zhenjie Ji, Yanjiao Shi, Qiangqiang Zhou
ICANN (2)4
2025 Frequency-guided Camouflaged Object Detection with Perceptual Enhancement and Dynamic Balance
abstract
Camouflaged object detection (COD) aims to segment concealed objects from their surroundings. However, most existing frequency-domain methods rely on simplistic fusion schemes with RGB features especially for objects with severe occlusion, varying scales, or ambiguous appearance. In this paper, we propose a frequency-guided COD network to explore how to utilize frequency-domain information to enhance the learning and presentation ability of RGB-domain features. Specifically, the frequency-aware query module is designed to selectively focus on essential RGB features by capturing and integrating long-term dependencies from frequency-domain information. Additionally, to enhance the model’s capability in identify objects of varying scales, the perception enhancement module is proposed to sufficiently integrate related and complementary features across adjacent levels. Finally, the dynamic balance module is introduced to aggregate multi-level features by adaptively balancing global contextual knowledge with local details information. Experiments on three widely-used benchmarks demonstrate the effectiveness and superiority of our network. The source code is available at https://github.com/iuueong/FPDNet.
Yuetong Li, Qing Zhang 0004, Qiangqiang Zhou, Yanjiao Shi
ICME4
2025 Dual-domain Collaboration Learning Network for Camouflaged Object Detection
abstract
Existing camouflaged object detection (COD) approaches primarily rely on RGB domain features to segment camouflaged objects. However, they face a major limitation: insensitivity to subtle differences in colors, textures and patterns, which contradicts the essence of COD that discerns imperceptible cues of camouflaged objects. To address this challenge, we propose a dual-domain collaborative learning network, which leverages frequency domain cues to collaborate with RGB domain features, therefore utilizing their unique and complementary strengths for effectively detecting camouflaged objects. Specifically, we propose the dual-domain feature integration (DFI) module, which aggregates cross-level RGB and frequency features to alleviate level-specific limitations, resulting in discriminative object features. Furthermore, we design the frequency-aware global localization (FGL) module, employing different frequency components to perceive global contexts. Additionally, we introduce the foreground-background separation learning (FSL) module to simultaneously capture semantic contexts and intricate details in both foreground and background. Extensive experimental results demonstrate the effectiveness of our network. Our results are available at https://github.com/ZhangQing0329/DCLNet
Jingming Wang, Qing Zhang 0004, Yanjiao Shi, Qiangqiang Zhou
IJCNN5
2025 HEFNet: Hierarchical Unimodal Enhancement and Multi-modal Fusion for RGB-T Salient Object Detection
abstract
RGB-Thermal salient object detection (RGB-T SOD) aims to identify and segment visually prominent objects by leveraging complementary information from RGB and thermal modalities. A key challenge lies in exploiting both the uniqueness and shared characteristics of these modalities to enhance their collaboration. Existing methods often ignore the optimization of unimodal features and the level-specific modality discrepancy, leading to noisy and redundant multi-modal feature representations. To address these limitations, we propose a novel RGB-T SOD network, HEFNet, which employs hierarchical unimodal enhancement and multi-modal fusion to achieve precise segmentation. Specifically, we introduce the unimodal feature enhancement (UFE) module, which refines RGB and thermal features by incorporating complementary information from adjacent levels, thereby enhancing saliency cues and suppressing noise distractions. Additionally, the hierarchical multi-modal fusion (HMF) module is designed to generate robust cross-modal feature representation. By employing tailored refinement and fusion strategies within the UFE and HMF modules, our network fully exploits the strengths of each modality, facilitating the generation of discriminative cross-modal features. Finally, the multi-level feature integration (MFI) module is introduced to progressively aggregate features across levels to ensure accurate saliency predictions. Extensive experiments demonstrate that our method achieves state-of-the-art performance, verifying its effectiveness and superiority over existing RGB-T SOD approaches. Our results are available at https://github.com/ZhangQing0329/HEFNet
Jiayun Wu, Qing Zhang 0004, Yanjiao Shi, Qiangqiang Zhou
IJCNN5
2025 FGNet: Feature Calibration and Guidance Refinement for Camouflaged Object Detection
abstract
A high-quality guidance cue is critical for accurately segmenting camouflaged objects from their visually similar surroundings. In this paper, we present a novel camouflaged object detection network to achieve complete segmentation predictions with fine-grained details by exploring how to generate and utilize the high-quality guidance cue. We first propose a feature self-calibration module to suppress noise and highlight camouflaged object regions from a contextual and spatial perspective. Based on the calibrated features, the proposed boundary-aware localization (BAL) module captures the coarse position information of camouflaged objects. Furthermore, the position information as the guidance cue is further refined iteratively by the context guidance refinement (CGR) module to effectively inform the network’s learning process. Finally, the progressive feature shrinking (PFS) module integrates adjacent features in a hierarchical manner to produce the final segmentation result. Experimental results on widely-used benchmark datasets demonstrate the effectiveness and superiority of our network. Our results are available at https://github.com/ZhangQing0329/FGNet
Qing Zhang 0004, Yanjiao Shi, Qiangqiang Zhou
IJCNN4
2025 Collaborative Perception and Dual-Stage Decoder Network for Camouflaged Object Detection
Zhenjie Ji, Yanjiao Shi, Qiangqiang Zhou
PRCV (18)4
2025 SAM2-LPNet: Saliency Guided and Laplacian Aware Fine-Tuning of SAM2 for Weakly Supervised Salient Object Detection
Yanjiao Shi, Qiangqiang Zhou
PRCV (3)4
2025 Mobility-Aware Cooperative Caching in IoVs Based on Secure Asynchronous Federated and Deep Reinforcement Learning
abstract
Edge content caching of Internet of Vehicles (IoVs) is a key technology for alleviating backhaul strain and reducing access latency. To protect the privacy of vehicular users, Federated learning (FL) is employed by sharing vehicles’ local models instead of data. However, vehicles may leave the coverage range of serving node before completing the local model training. To enhance model aggregation efficiency, asynchronous Federated learning (AFL) is employed, which allows asynchronous aggregation without waiting for all vehicles to update their local models. In practice, the local models are susceptible to malicious tampering during the global aggregation process. To solve this problem, we propose a secure AFL (SAFL) framework by incorporating a Z-score-based weight detection method within AFL. Moreover, To improve caching efficiency and adapt to the highly dynamic IoV environments, we introduce an innovative proactive caching approach by combining a conditional variational autoencoder and generative adversarial network to predict popular contents, thereby improving the cache hit ratio. Additionally, based on the prediction results of popular content, we optimize intelligent decision-making using multiagent deep reinforcement learning (DRL) to reduce the content transmission delay. Extensive simulations are performed based on real-world datasets and experimental results demonstrate that the proposed SAFL and multiagent DRL hybrid technique outperforms other baseline approaches.
Xuefang Nie, Chen Wang 0069, Tianqing Zhou, Qiangqiang Zhou, Xusheng Zhu, Jiliang Zhang 0001
IEEE Internet Things J.4
2025 Semantic-Orthogonal Multi-modal Attention Network for RGB-D Salient Object Detection
Jiawei Xu 0007, Qiangqiang Zhou, Jiacong Yu, Chen Liao
Vis. Comput.2
2024 Transformer-Based Depth Optimization Network for RGB-D Salient Object Detection
Yanjiao Shi, Qiangqiang Zhou, Liu Cui
ICPR (21)4
2024 Unified Audio-Visual Saliency Model for Omnidirectional Videos With Spatial Audio
abstract
Spatial audio is a crucial component of omnidirectional videos (ODVs), which can provide an immersive experience by enabling viewers to perceive sound sources in all directions. However, most visual attention modeling works for ODVs focus only on visual cues, and audio modality is rather rarely considered. Additionally, the existing audio-visual saliency models for ODVs lack spatial audio location-awareness (i.e. sound source location-agnostic) and audio content attributes discriminability (i.e. audio content attributes-agnostic). To this end, we propose a novel audio-visual perception saliency (AVPS) model with spatial audio location-awareness and audio content attributes-adaptive to efficiently address the problem of fixation prediction in ODVs. Specifically, we first utilize the improved group equivariant convolutional neural network (G-CNN) with eidetic 3D LSTM (E3D-LSTM) to extract spatial-temporal visual features. Then we perceive sound source locations by computing the audio energy map (AEM) of the audio information in ODVs. Subsequently, we introduce SoundNet to extract audio features with multiple attributes. Finally, we develop an audio-visual feature fusion module to adaptively integrate spatial-temporal visual features and spatial auditory information to generate the final audio-visual saliency map. Extensive experiments in three audio modalities validate the effectiveness of the proposed model. Meanwhile, the performance of the proposed model is superior to the other 10 state-of-the-art saliency models.
Dandan Zhu 0001, Kaiwei Zhang, Qiangqiang Zhou, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001
IEEE Trans. Multim.4
2023 Human attention based movie summarization: Dataset and baseline model
abstract
A movie summarization model can automatically edit a condensed version of a movie by selecting keyframes . Some previous works have proposed some movie summarizers based on traditional methods or recent neural networks and achieved some progress. Despite the demonstrated successes, there are some limitations: (1) previous works mainly resort to hand-crafted heuristics and most of them are unsupervised; (2) currently there is no publicly suitable dataset available for the supervised movie summarization; (3) existing works only focus on the movies themselves while neglecting the audiences, who have the most to say in which part of the movie is more attractive. To break through the aforementioned limitations, we establish a movie summarization dataset Movie50 and propose a novel human attention based annotation pipeline. Furthermore, we propose the A/V-MSNet, an audiovisual neural network that takes advantage of spatio-temporal visual and auditory information to better simulate human attention as well as exploit more plentiful information. The network is designed, trained end-to-end, and evaluated on the public dataset and our dataset. Extensive experiments demonstrate the superiority of the proposed method.
Defang Zhao, Dandan Zhu 0001, Xiongkuo Min, Jiaomin Yue, Kaiwei Zhang, Qiangqiang Zhou, Guangtao Zhai, Xiaokang Yang 0001
Neurocomputing6
2023 Decoupled dynamic group equivariant filter for saliency prediction on omnidirectional image
Dandan Zhu 0001, Kaiwei Zhang, Qiangqiang Zhou, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001
Neurocomputing4
2023 A Novel Lightweight Audio-visual Saliency Model for Videos
abstract
Audio information has not been considered an important factor in visual attention models regardless of many psychological studies that have shown the importance of audio information in the human visual perception system. Since existing visual attention models only utilize visual information, their performance is limited but also requires high-computational complexity due to the limited information available. To overcome these problems, we propose a lightweight audio-visual saliency (LAVS) model for video sequences. To the best of our knowledge, this article is the first trial to utilize audio cues for an efficient deep-learning model for the video saliency estimation. First, spatial-temporal visual features are extracted by the lightweight receptive field block (RFB) with the bidirectional ConvLSTM units. Then, audio features are extracted by using an improved lightweight environment sound classification model. Subsequently, deep canonical correlation analysis (DCCA) aims at capturing the correspondence between audio and spatial-temporal visual features, thus obtaining a spatial-temporal auditory saliency. Lastly, the spatial-temporal visual and auditory saliency are fused to obtain the audio-visual saliency map. Extensive comparative experiments and ablation studies validate the performance of the LAVS model in terms of effectiveness and complexity.
Dandan Zhu 0001, Xuan Shao, Qiangqiang Zhou, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Human Attention Based Movie Summarization: Dataset and Baseline Model
abstract
The movie summarization model can automatically edit a condensed and succinct version of the movie by selecting the keyframes. Previous works mainly resort to hand-crafted heuristics and most of them are unsupervised. Supervised movie summarization is a new research field and, there is currently no publicly suitable dataset available. Moreover, existing works only focus on the movies themselves while neglecting the audiences, who have the most say in which part of the movie is more attractive. To deal with the aforementioned limitations, we establish a human attention based movie summarization dataset Movie50. Specifically, we explore the human attention variations when watching videos and have the following findings: (1) The attention of humans is concentrated when watching keyframes. (2) The attention of humans is distracted when watching non-keyframes. Inspired by these findings, we collect the eye fixations of 20 participants when watching 50 movies and propose a novel human attention based annotation pipeline. In addition, we introduce A/V-MSNet, an audiovisual neural network that takes advantage of spatio-temporal visual and auditory information to better model human attention as well as exploit more plentiful information. Extensive experiments demonstrate the superiority of the proposed method.
Defang Zhao, Dandan Zhu 0001, Xiongkuo Min, Jiaomin Yue, Kaiwei Zhang, Qiangqiang Zhou, Guangtao Zhai, Xiaokang Yang 0001
ICME6
2021 A Lightweight Saliency Prediction Model for Omnidirectional Images
abstract
At present, most high-performing saliency prediction models for omnidirectional images (ODIs) depend on deeper or wider convolutional neural networks (CNNs), benefiting from their superior feature representation capability but suffering from high computational costs. To address this issue, we propose a novel lightweight saliency prediction model to predict the eye fixations on ODIs. Specifically, our proposed model consists of three modules: a lightweight feature representation module, a supervised attention module, and a dynamic convolution aggregation module. Different from the existing saliency prediction models, our proposed model is the first to introduce the dynamic convolution into the saliency prediction and aggregate multiple parallel convolution kernels dynamically based on their attention. Such a dynamic convolution operation is not only computationally efficient (small kernel size), but also increases the feature representation capability since these convolution kernels are aggregated in a non-linear manner via attention. Experimental results on two benchmark datasets show that our model is lightweight and outperforms other state-of-the-art methods.
Dandan Zhu 0001, Yongqing Chen, Defang Zhao, Xiongkuo Min, Qiangqiang Zhou, Shaobo Yu, Guangtao Zhai, Xiaokang Yang 0001
ICME5
2021 Lavs: A Lightweight Audio-Visual Saliency Prediction Model
abstract
Audio information is essential for guiding human attention and visual perception, which has been verified by many comprehensive psychological studies. However, the audio modality has been rather neglected in modeling visual attention, most of the current visual attention models heavily depend on visual information. Additionally, current existing high-performing visual attention models rely on deeper convolution neural networks (CNNs), benefiting from their extraordinary feature learning ability but incurring high computational cost. To this end, we propose a novel lightweight audio-visual saliency (LAVS) model to efficiently address the problem of fixation prediction in videos. To the best of our knowledge, our proposed model constitutes the first attempt to exploit a lightweight network and combines the visual and audio cues to perform saliency estimation in videos. Specifically, our proposed model consists of four modules, which are spatial-temporal visual saliency estimation module, audio features extraction module, source sound localization module, and audio-visual saliency fusion module. Extensive experiments across datasets validate the effectiveness and real-time performance of the proposed LAVS model, which outperforms the other state-of-the-art methods.
Dandan Zhu 0001, Defang Zhao, Xiongkuo Min, Tian Han 0001, Qiangqiang Zhou, Shaobo Yu, Yongqing Chen, Guangtao Zhai, Xiaokang Yang 0001
ICME5
2021 Saliency prediction on omnidirectional images with attention-aware feature fusion network
Yongqing Chen, Defang Zhao, Qiangqiang Zhou, Xiaokang Yang 0009
Appl. Intell.4
2021 RANSP: Ranking attention network for saliency prediction on omnidirectional images
Dandan Zhu 0001, Yongqing Chen, Xiongkuo Min, Yucheng Zhu, Qiangqiang Zhou, Guangtao Zhai, Xiaokang Yang 0001
Neurocomputing6
2021 Towards multi-scale deep features learning with correlation metric for person re-identification
Dandan Zhu 0001, Qiangqiang Zhou, Tian Han 0001, Yongqing Chen, Defang Zhao, Xiaokang Yang 0001
Knowl. Based Syst.2
2020 Ransp: Ranking Attention Network For Saliency Prediction On Omnidirectional Images
abstract
Various convolutional neural network (CNN)-based methods have shown the ability to boost the performance of saliency prediction on omnidirectional images (ODIs). However, these methods are limited by sub-optimal accuracy, because not all the features extracted by the CNN model are not useful for the final fine-grained saliency prediction. Features are redundant and have negative impact on the final fine-grained saliency prediction. To tackle this problem, we propose a novel Ranking Attention Network for saliency prediction (RANSP) of head fixations on ODIs. Specifically, the part-guided attention (PA) module and channel-wise feature (CF) extraction module are integrated in a unified framework and are trained in an end-to-end manner for fine-grained saliency prediction. To better utilize the channel-wise feature map, we further propose a new Ranking Attention Module (RAM), which automatically ranks and selects these maps based on scores for fine-grained saliency prediction. Extensive experiments are conducted to show the effectiveness of our method for saliency prediction of ODIs.
Dandan Zhu 0001, Yongqing Chen, Tian Han 0001, Defang Zhao, Yucheng Zhu, Qiangqiang Zhou, Guangtao Zhai, Xiaokang Yang 0001
ICME6
2020 Saliency Prediction on Omnidirectional Images with Brain-Like Shallow Neural Network
abstract
Deep feedforward convolutional neural networks (CNNs) perform well in the saliency prediction of omnidirectional images (ODIs), and have become the leading class of candidate models of the visual processing mechanism in the primate ventral stream. These CNNs have evolved from shallow network architecture to extremely deep and branching architecture to achieve superb performance in various vision tasks, yet it is unclear how brain-like they are. In particular, these deep feedforward CNNs are difficult to mapping to ventral stream structure of the brain visual system due to their vast number of layers and missing biologically-important connections, such as recurrence. To tackle this issue, some brain-like shallow neural networks are introduced. In this paper, we propose a novel brain-like network model for saliency prediction of head fixations on ODIs. Specifically, our proposed model consists of three modules: a CORnet-S module, a template feature extraction module and a ranking attention module (RAM). The CORnet-S module is a lightweight artificial neural network (ANN) with four anatomically mapped areas (V1, V2, V4 and IT) and it can simulate the visual processing mechanism of ventral visual stream in the human brain. The template features extraction module is introduced to extract attention maps of ODIs and provide guidance for the feature ranking in the following RAM module. The RAM module is used to rank and select features that are important for fine-grained saliency prediction. Extensive experiments have validated the effectiveness of the proposed model in predicting saliency maps of ODIs, and the proposed model outperforms other state-of-the-art methods with similar scale.
Dandan Zhu 0001, Yongqing Chen, Xiongkuo Min, Defang Zhao, Yucheng Zhu, Qiangqiang Zhou, Xiaokang Yang 0001, Tian Han 0001
ICPR6
2018 Salient object detection via a local and global method based on deep residual network
Dandan Zhu 0001, Ye Luo 0004, Xuan Shao, Qiangqiang Zhou, Laurent Itti
J. Vis. Commun. Image Represent.5
2017 Robust Non-Rigid Point Set Registration Using Spatially Constrained Gaussian Fields
abstract
Estimating transformations from degraded point sets is necessary for many computer vision and pattern recognition applications. In this paper, we propose a robust non-rigid point set registration method based on spatially constrained context-aware Gaussian fields. We first construct a context-aware representation (e.g., shape context) for assignment initialization. Then, we use a graph Laplacian regularized Gaussian fields to estimate the underlying transformation from the likely correspondences. On the one hand, the intrinsic manifold is considered and used to preserve the geometrical structure, and a priori knowledge of the point set is extracted. On the other hand, by using the deterministic annealing, the presented method is extended to a projected high-dimensional feature space, i.e., reproducing kernel Hilbert space through a kernel trick to solve the transformation, in which the local structure is propagated by the coarse-to-fine scaling strategy. In this way, the proposed method gradually recovers much more correct correspondences, and then estimates the transformation parameters accurately and robustly when facing degradations. Experimental results on 2D and 3D synthetic and real data (point sets) demonstrate that the proposed method reaches better performance than the state-of-the-art algorithms.
Gang Wang 0008, Qiangqiang Zhou, Yufei Chen 0002
IEEE Trans. Image Process.2
2016 Context-Aware Gaussian Fields for Non-rigid Point Set Registration
abstract
Point set registration (PSR) is a fundamental problem in computer vision and pattern recognition, and it has been successfully applied to many applications. Although widely used, existing PSR methods cannot align point sets robustly under degradations, such as deformation, noise, occlusion, outlier, rotation, and multi-view changes. This paper proposes context-aware Gaussian fields (CA-LapGF) for nonrigid PSR subject to global rigid and local non-rigid geometric constraints, where a laplacian regularized term is added to preserve the intrinsic geometry of the transformed set. CA-LapGF uses a robust objective function and the quasi-Newton algorithm to estimate the likely correspondences, and the non-rigid transformation parameters between two point sets iteratively. The CA-LapGF can estimate non-rigid transformations, which are mapped to reproducing kernel Hilbert spaces, accurately and robustly in the presence of degradations. Experimental results on synthetic and real images reveal that how CA-LapGF outperforms state-of-the-art algorithms for non-rigid PSR.
Gang Wang 0008, Zhicheng Wang 0022, Yufei Chen 0002, Qiangqiang Zhou
CVPR4
2016 Removing mismatches for retinal image registration via multi-attribute-driven regularized mixture model
Gang Wang 0008, Zhicheng Wang 0022, Yufei Chen 0002, Qiangqiang Zhou
Inf. Sci.4
2015 Contour-Based Plant Leaf Image Segmentation Using Visual Saliency
Qiangqiang Zhou, Zhicheng Wang 0022, Yufei Chen 0002
ICIG (2)1
2014 Robust Point Matching Using Mixture of Asymmetric Gaussians for Nonrigid Transformation
Gang Wang 0008, Zhicheng Wang 0022, Qiangqiang Zhou
ACCV (4)4