Qing Zhang 0004

dblp:68/1429-4 · DBLP profile ↗
← Back
42ranked-venue papers
14as first author
33since 2021 · last 2025
0000-0001-5318-995XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 8 first-author · 17 since 2021Artificial intelligence and machine learning · 20 · 6 first-author · 16 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Rethinking Camouflaged Object Detection via Foreground-Background Interactive Learning
abstract
Camouflaged object detection focuses on the challenge of segmenting objects that visually blend into their background. The effectiveness of camouflage strategies hinges on how well objects interact with their background to minimize their visibility. Based on this insight, we propose a novel Foreground-Background Interactive Learning Network (FBINet), which independently decouples foreground and background information, and facilitates bi-directional interactions. This allows the network to progressively complement each other, leading to high-quality predictions with clear boundaries. To the best of our knowledge, this is the first attempt to tackle the camouflaged object detection task through interactive learning between foreground and background, which not only better reveals camouflage patterns but also offers a new perspective in this field. Experimental results on three datasets demonstrate that our proposed FBINet outperforms current state-of-the-art methods in performance while maintaining a low computational cost, making it applicable for real-world scenarios. The code will be available at https://github.com/bbdjj/FBINet.
Qing Zhang 0004, Jiayun Wu
ICASSP2
2025 Camouflaged Object Detection with CNN-Transformer Harmonization and Calibration
abstract
Camouflaged object detection (COD) aims to segment objects that visually blend into their surroundings. However, the subtle differences between camouflaged objects and the background make this task highly challenging. Therefore, how to represent and learn local details and global contexts is crucial for improving detection performance. In this paper, we propose a novel COD network which synergistically leverages the distinct but complementary local and global knowledge to capture the camouflaged objects and identify imperceptible boundaries. Specifically, we design a Feature Coherence Harmonization module to integrate intra-layer features by bridging the knowledge gap between convolutional neural network (CNN) features, which focus on local patterns, and Transformer features, which capture global relationships. Furthermore, we propose a Cross-layer Feature Calibration Module that adaptively aligns inter-layer features, progressively aggregating diverse information to achieve an accurate prediction. Experimental results on COD benchmark datasets demonstrate that the proposed network significantly outperforms state-of-the-art approaches.
Qing Zhang 0004, Yuetong Li
ICASSP2
2025 Frequency-guided Camouflaged Object Detection with Perceptual Enhancement and Dynamic Balance
abstract
Camouflaged object detection (COD) aims to segment concealed objects from their surroundings. However, most existing frequency-domain methods rely on simplistic fusion schemes with RGB features especially for objects with severe occlusion, varying scales, or ambiguous appearance. In this paper, we propose a frequency-guided COD network to explore how to utilize frequency-domain information to enhance the learning and presentation ability of RGB-domain features. Specifically, the frequency-aware query module is designed to selectively focus on essential RGB features by capturing and integrating long-term dependencies from frequency-domain information. Additionally, to enhance the model’s capability in identify objects of varying scales, the perception enhancement module is proposed to sufficiently integrate related and complementary features across adjacent levels. Finally, the dynamic balance module is introduced to aggregate multi-level features by adaptively balancing global contextual knowledge with local details information. Experiments on three widely-used benchmarks demonstrate the effectiveness and superiority of our network. The source code is available at https://github.com/iuueong/FPDNet.
Yuetong Li, Qing Zhang 0004, Qiangqiang Zhou, Yanjiao Shi
ICME3
2025 Dual-domain Collaboration Learning Network for Camouflaged Object Detection
abstract
Existing camouflaged object detection (COD) approaches primarily rely on RGB domain features to segment camouflaged objects. However, they face a major limitation: insensitivity to subtle differences in colors, textures and patterns, which contradicts the essence of COD that discerns imperceptible cues of camouflaged objects. To address this challenge, we propose a dual-domain collaborative learning network, which leverages frequency domain cues to collaborate with RGB domain features, therefore utilizing their unique and complementary strengths for effectively detecting camouflaged objects. Specifically, we propose the dual-domain feature integration (DFI) module, which aggregates cross-level RGB and frequency features to alleviate level-specific limitations, resulting in discriminative object features. Furthermore, we design the frequency-aware global localization (FGL) module, employing different frequency components to perceive global contexts. Additionally, we introduce the foreground-background separation learning (FSL) module to simultaneously capture semantic contexts and intricate details in both foreground and background. Extensive experimental results demonstrate the effectiveness of our network. Our results are available at https://github.com/ZhangQing0329/DCLNet
Jingming Wang, Qing Zhang 0004, Yanjiao Shi, Qiangqiang Zhou
IJCNN2
2025 HEFNet: Hierarchical Unimodal Enhancement and Multi-modal Fusion for RGB-T Salient Object Detection
abstract
RGB-Thermal salient object detection (RGB-T SOD) aims to identify and segment visually prominent objects by leveraging complementary information from RGB and thermal modalities. A key challenge lies in exploiting both the uniqueness and shared characteristics of these modalities to enhance their collaboration. Existing methods often ignore the optimization of unimodal features and the level-specific modality discrepancy, leading to noisy and redundant multi-modal feature representations. To address these limitations, we propose a novel RGB-T SOD network, HEFNet, which employs hierarchical unimodal enhancement and multi-modal fusion to achieve precise segmentation. Specifically, we introduce the unimodal feature enhancement (UFE) module, which refines RGB and thermal features by incorporating complementary information from adjacent levels, thereby enhancing saliency cues and suppressing noise distractions. Additionally, the hierarchical multi-modal fusion (HMF) module is designed to generate robust cross-modal feature representation. By employing tailored refinement and fusion strategies within the UFE and HMF modules, our network fully exploits the strengths of each modality, facilitating the generation of discriminative cross-modal features. Finally, the multi-level feature integration (MFI) module is introduced to progressively aggregate features across levels to ensure accurate saliency predictions. Extensive experiments demonstrate that our method achieves state-of-the-art performance, verifying its effectiveness and superiority over existing RGB-T SOD approaches. Our results are available at https://github.com/ZhangQing0329/HEFNet
Jiayun Wu, Qing Zhang 0004, Yanjiao Shi, Qiangqiang Zhou
IJCNN2
2025 FGNet: Feature Calibration and Guidance Refinement for Camouflaged Object Detection
abstract
A high-quality guidance cue is critical for accurately segmenting camouflaged objects from their visually similar surroundings. In this paper, we present a novel camouflaged object detection network to achieve complete segmentation predictions with fine-grained details by exploring how to generate and utilize the high-quality guidance cue. We first propose a feature self-calibration module to suppress noise and highlight camouflaged object regions from a contextual and spatial perspective. Based on the calibrated features, the proposed boundary-aware localization (BAL) module captures the coarse position information of camouflaged objects. Furthermore, the position information as the guidance cue is further refined iteratively by the context guidance refinement (CGR) module to effectively inform the network’s learning process. Finally, the progressive feature shrinking (PFS) module integrates adjacent features in a hierarchical manner to produce the final segmentation result. Experimental results on widely-used benchmark datasets demonstrate the effectiveness and superiority of our network. Our results are available at https://github.com/ZhangQing0329/FGNet
Qing Zhang 0004, Yanjiao Shi, Qiangqiang Zhou
IJCNN2
2025 Rethinking Lightweight and Efficient Human Pose Estimation with Star Operation Reconstruction
Zhoujie Xu, Qing Zhang 0004, Huawen Liu
KSEM (4)3
2025 CGCOD: Class-Guided Camouflaged Object Detection
abstract
Camouflaged Object Detection (COD) aims to identify objects that blend seamlessly into their surroundings. The inherent visual complexity of camouflaged objects, including their low contrast with the background, diverse textures, and subtle appearance variations, often obscures semantic cues, making accurate segmentation highly challenging. Existing methods primarily rely on visual features, which are insufficient to handle the variability and intricacy of camouflaged objects, leading to unstable object perception capability and ambiguous segmentation results. To tackle these limitations, we introduce a novel COD task, class-guided camouflaged object detection (CGCOD), which extends the conventional COD task by incorporating object-specific class knowledge to enhance detection robustness and accuracy. To facilitate this task, we present a new dataset, CamoClass, comprising camouflaged objects with class annotations. Furthermore, we propose a multi-stage framework, CGNet, which incorporates a plug-and-play class prompt generator and a simple yet effective class-guided detector. This establishes a new paradigm for COD, bridging the gap between contextual understanding and class-guided detection. Extensive experimental results demonstrate the effectiveness of our flexible framework in improving the performance of proposed and existing detectors by leveraging class-level textual information. The Camoclass dataset and the corresponding source code will be made publicly available upon acceptance at: https://github.com/bbdjj/CGCOD.
Qing Zhang 0004, Jiayun Wu, Youwei Pang
ACM Multimedia2
2025 FINet: Detecting Camouflaged Objects Using Frequency-Aware Information
Chengfeng Zhu, Zeming Liu, Qing Zhang 0004, Xuefeng Wu
PRCV (18)3
2025 HRPVT: High-Resolution Pyramid Vision Transformer for medium and small-scale human pose estimation
Zhoujie Xu, Qing Zhang 0004, Xiaodi Jiang
Neurocomputing3
2025 FLRNet: A bio-inspired three-stage network for Camouflaged Object Detection via filtering, localization and refinement
Qing Zhang 0004, Yuetong Li
Neurocomputing2
2025 Boundary-and-object collaborative learning network for camouflaged object detection
Chenyu Zhuang, Qing Zhang 0004, Xinxin Yuan
Image Vis. Comput.2
2025 Bilateral decoupling complementarity learning network for camouflaged object detection
Yuetong Li, Qing Zhang 0004
Knowl. Based Syst.3
2024 Bi-directional Boundary-object interaction and refinement network for Camouflaged Object Detection
abstract
Due to the high intrinsic similarity between the camouflaged objects and the background, the predicted edge cue might be inaccurate or even erroneous. This will degrade the detection performance when such edge cues are directly integrated with the camouflaged object features. To solve this issue, we propose a bi-directional camouflaged object detection network to progressive complement and correct the boundary feature and the object feature in an interactive learning manner, thereby achieving prediction with fine structures. Specifically, we design a bi-directional transmission and correction (BTC) module, which contains the boundary and object interaction (BOI) module and the Feature Calibration Fusion (FCF) module, to explore the correlation between the boundary feature and the object feature and then provide the important complementary cues to refine each other, thereby progressively compensating the deficiencies and correcting the mistakes. Experimental results on four benchmark datasets demonstrate the effectiveness and superiority of our network. Our code is publicly available at: https://github.com/Jcogito/BIRNet.
Jicheng Yang, Qing Zhang 0004, Yuetong Li, Zeming Liu
ICME2
2024 Detecting camouflaged objects via cross-level context supplement
Qing Zhang 0004, Weiqi Yan 0002, Yanjiao Shi
Appl. Intell.1
2024 BMFNet: Bifurcated multi-modal fusion network for RGB-D salient object detection
Chenwang Sun, Qing Zhang 0004, Chenyu Zhuang, Mingqian Zhang
Image Vis. Comput.2
2024 AFINet: Camouflaged object detection via Attention Fusion and Interaction Network
Qing Zhang 0004, Weiqi Yan 0002
J. Vis. Commun. Image Represent.1
2024 Multi-branch feature fusion and refinement network for salient object detection
Yanjiao Shi, Qianqian Guo, Qing Zhang 0004, Liu Cui
Multim. Syst.5
2024 Salient Object Detection With Edge-Guided Learning and Specific Aggregation
abstract
Recently, the performance of salient object detection (SOD) has been significantly improved by utilizing edge information for auxiliary training. However, the extraction and utilization of edge cues and multi-level feature fusion are still two issues in existing edge-aware models. In this paper, we devise a novel SOD network with edge-guided learning and specific aggregation, named ELSA-Net, to cooperatively address these two issues. First, we propose the edge-guided learning strategy, which utilizes edge cues as low-level guidance to improve saliency prediction. Specifically, we design a two-stream model that uses a saliency branch and an edge branch to detect the interior and the boundary of salient objects, respectively. Then, an edge-guided interaction module (EGI) is further designed to achieve feature enhancement by embedding edge information into the saliency branch as the spatial weights. In addition, two specific aggregation modules are proposed for the progressive fusion of multi-level features in the above two streams, thus making full use of semantic and detailed information. The high-level interactive fusion module (HIF) leverages the correlation between two deeper features to obtain more powerful global contexts. And the low-level weighted fusion module (LWF) focuses on the complement of fine information by selectively integrating input features. Extensive experiments show that the proposed approach outperforms 19 state-of-the-art methods on five datasets, which validates its effectiveness both quantitatively and qualitatively.
Qing Zhang 0004
IEEE Trans. Circuits Syst. Video Technol.2
2023 CFANet: A Cross-layer Feature Aggregation Network for Camouflaged Object Detection
abstract
Deep learning-based camouflaged object detection approaches have achieved great progress in recent years. However, it is still challenging to accurately identify the camouflaged objects from the highly similar backgrounds. In this paper, we design a novel cross-layer feature aggregation network (CFANet) for camouflaged object detection to explore how to effectively aggregate multi-level and multi-scale features generated from the backbone network by excavating the similarities and differences of features at different levels. Firstly, we design a cross-layer feature fusion (CLFF) module to generate discriminative features by fusing and refining the multi-level side-output features with similar characteristics in the same feature group. Secondly, we design a uniqueness enhancement (UE) strategy to respectively emphasize the superiority of deep features and shallow features in locating the camouflaged objects and sharpening the structure details. Extensive experiments on four benchmarks are conducted to demonstrate that the proposed CFANet network performs favorably against 11 state-of-the-art camouflaged object detection methods, demonstrating the effectiveness and superiority of our method. The code and results can be found from the link of https://github.com/ZhangQing0329/CFANet
Qing Zhang 0004, Weiqi Yan 0002
ICME1
2023 SC2Net: Scale-aware Crowd Counting Network with Pyramid Dilated Convolution
Lanjun Liang, Huailin Zhao, Fangbo Zhou, Qing Zhang 0004, Zhili Song, Qingxuan Shi
Appl. Intell.4
2023 Depth cue enhancement and guidance network for RGB-D salient object detection
Xiang Li 0090, Qing Zhang 0004, Weiqi Yan 0002
J. Vis. Commun. Image Represent.2
2023 Polyp-Mixer: An Efficient Context-Aware MLP-Based Paradigm for Polyp Segmentation
abstract
Precise and efficient polyp segmentation plays a crucial role in colonoscopy, which is important for the prevention of colorectal cancer. Despite CNN-based methods have achieved great progress in the polyp segmentation task, they are incapable of modeling long-range dependencies. Transformer-based models utilize self-attention mechanism to overcome this problem while suffering from heavy computing cost. Benefiting from simple structures, MLP-based models seem to be an alternative. However, they struggle with dealing with flexible input scales and modeling long-term dependencies. Both of these two factors are important for image segmentation, which could explain why the MLP architecture performs poorly compared to the Transformer. To remedy this issue, we propose a novel Polyp-Mixer, which utilizes MLP-based structures in both encoder and decoder. In particular, we use CycleMLP as the encoder to overcome the fixed input scale issue. Besides, we propose a Multi-head Mixer by converting the current CycleMLP into a Multi-head fashion, allowing our model to explore rich context information from various subspaces. In addition, we build a powerful Contextual Bridger Module between the encoder and decoder, which can capture semantics from larger receptive fields and combine them with various decoder layers. Experiments demonstrate the proposed method with fewer parameters ($\sim 16\text{M}$) achieves SOTA on 4 public benchmarks. Our code will be released athttps://github.com/shijinghuihub/Polyp-Mixer
Jing-Hui Shi, Qing Zhang 0004, Yuhao Tang, Zhong-Qun Zhang
IEEE Trans. Circuits Syst. Video Technol.2
2023 TCRNet: A Trifurcated Cascaded Refinement Network for Salient Object Detection
abstract
Due to the rapid development of deep learning, salient object detection has achieved promising results in recent years. However, how to construct a powerful saliency detection network to generate saliency maps with completely highlighted salient objects and effectively suppressed background noise still remains an open and challenging problem. In this paper, we propose a novel trifurcated cascaded refinement network (TCRNet) to explore multi-level feature fusion and global information representation. Specifically, all the side-output features from the backbone network are divided into three branches (i.e., global information, deep information and shallow information) according to feature similarities, and then the global contexts, semantic knowledge and fined details respectively generated by the three branches are integrated in a cascaded refinement way. This refinement strategy aims to guide the feature learning of important regions and suppress background noises. Moreover, we propose a cross-level feature fusion (CFF) module to fuse multi-level similar features within a same branch by exploring their interrelationship, thus capturing the important semantic information and structure details. To this end, the complementarity of similar features is utilized to refine feature at each level by sharing an encoded spatial weight map. Besides, we design a context-aware feature learning (CFL) module to capture global contexts in a single layer by taking account of both single-feature representation and multi-feature interaction, thus providing important location information of salient objects. Extensive experiments on five benchmarks are conducted to demonstrate that the proposed network TCRNet performs favorably against 20 state-of-the-art salient object detection methods on five benchmark datasets, indicating the effectiveness and superiority of our method.
Qing Zhang 0004
IEEE Trans. Circuits Syst. Video Technol.1
2022 Residual attentive feature learning network for salient object detection
Qing Zhang 0004, Yanjiao Shi
Neurocomputing1
2022 R2Net: Residual refinement network for salient object detection
Qiuwei Liang, Qianqian Guo, Qing Zhang 0004, Yanjiao Shi
Image Vis. Comput.5
2022 Attention guided contextual feature fusion network for salient object detection
Yanjiao Shi, Qing Zhang 0004, Liu Cui, Yugen Yi
Image Vis. Comput.3
2022 COMAL: compositional multi-scale feature enhanced learning for crowd counting
Fangbo Zhou, Huailin Zhao, Qing Zhang 0004, Lanjun Liang, Zuodong Duan
Multim. Tools Appl.4
2022 Progressive Dual-Attention Residual Network for Salient Object Detection
abstract
Due to the rapid development of deep learning, the performance of salient object detection has been constantly refreshed. Nevertheless, it is still challenging for existing methods to distinguish the location of salient objects and retain fine structural details. In this paper, a novel progressive dual-attention residual network (PDRNet) is proposed to exploit two complementary attention maps to guide residual learning, thus progressively refining prediction in a coarse-to-fine manner. We design a dual-attention residual module (DRM) to achieve residual refinement with the help of the dual attention (DA) scheme. Specifically, an attention map and its corresponding reverse attention map are used to make the network be aware of learning residual details from the perspective of the salient and non-salient regions, thus utilizing their complementarity to correct the mistakes of object parts and boundary details. Besides, a hierarchical feature screening module (HFSM) is designed to capture more powerful global contextual knowledge for locating salient objects. It establishes cross-scale skip connections among multi-scale features and utilizes the intra-channel dependency of these scales to enhance information interaction and feature representation. Extensive experiments have proved that our proposed PDRNet performs favorably against 18 state-of-the-art competitors on five benchmark datasets, demonstrating the effectiveness and superiority of our method.
Qing Zhang 0004
IEEE Trans. Circuits Syst. Video Technol.2
2021 MSCANet: Adaptive Multi-scale Context Aggregation Network for Congested Crowd Counting
Huailin Zhao, Fangbo Zhou, Qing Zhang 0004, Yanjiao Shi, Lanjun Liang
MMM (2)4
2021 Edge-aware salient object detection network via context guidance
Qing Zhang 0004
Image Vis. Comput.2
2021 Salient object detection network with multi-scale feature refinement and boundary feedback
Qing Zhang 0004, Xiang Li 0090
Image Vis. Comput.1
2021 Global and local information aggregation network for edge-aware salient object detection
Qing Zhang 0004, Yanjiao Shi, Jiajun Lin
J. Vis. Commun. Image Represent.1
2020 Attentive feature integration network for detecting salient objects in images
Qing Zhang 0004, Wenzhao Cui, Yanjiao Shi
Neurocomputing1
2020 Attention and boundary guided salient object detection
Qing Zhang 0004, Yanjiao Shi
Pattern Recognit.1
2019 Overview on Vision-Based 3D Object Recognition Methods
Tianzhen Dong, Qing Zhang 0004, Wenju Li, Liang Xiong
ICIG (2)3
2019 Hierarchical Salient Object Detection Network with Dense Connections
Qing Zhang 0004, Jianchen Shi 0002, Baochuan Zuo, Tianzhen Dong
ICIG (1)1
2019 Multi-level and multi-scale deep saliency network for salient object detection
Qing Zhang 0004, Jiajun Lin, Jingjing Zhuge, Wenhao Yuan 0001
J. Vis. Commun. Image Represent.1
2018 Salient object detection via compactness and objectness cues
Qing Zhang 0004, Jiajun Lin, Wenju Li, Yanjiao Shi, Guogang Cao
Vis. Comput.1
2017 Two-stage absorbing Markov chain for salient object detection
abstract
We propose a simple but effective approach to detect salient objects by exploring both patch-level and object-level cues under the framework of absorbing Markov chain. Saliency detection is carried out in a two-stage scheme. In the first stage, we conduct random walk on absorbing Markov chain with coarsely selected background seeds in the boundary. The result is integrated with a objectness map which is generated by finding potential object candidates to boost the saliency detection for object completeness. And in the second stage, we use the refined background seeds computed by the first stage as absorbing nodes for the absorbing Markov chain to obtain the final saliency map. Experimental results on four publicly available datasets demonstrate the robustness and efficiency of our proposed approach against 8 state-of-the-art methods in terms of five performance criterions.
Qing Zhang 0004, Desi Luo, Wenju Li, Yanjiao Shi, Jiajun Lin
ICIP1
2017 Salient object detection via color and texture cues
Qing Zhang 0004, Jiajun Lin, Yanyun Tao, Wenju Li, Yanjiao Shi
Neurocomputing1
2016 A systematic EHW approach to the evolutionary design of sequential circuits
Yanyun Tao, Qing Zhang 0004
Soft Comput.2