Yongqiang Zhang 0007

dblp:67/5744-7 · DBLP profile ↗
← Back
31ranked-venue papers
10as first author
22since 2021 · last 2026
0000-0002-0437-7337ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 9 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Causal-Tune: Mining Causal Factors from Vision Foundation Models for Domain Generalized Semantic Segmentation
abstract
Fine-tuning Vision Foundation Models (VFMs) with a small number of parameters has shown remarkable performance in Domain Generalized Semantic Segmentation (DGSS). Most existing works either train lightweight adapters or refine intermediate features to achieve better generalization on unseen domains. However, they both overlook the fact that long-term pre-trained VFMs often exhibit artifacts, which hinder the utilization of valuable representations and ultimately degrade DGSS performance. Inspired by causal mechanisms, we observe that these artifacts are associated with non-causal factors, which usually reside in the low- and high-frequency components of the VFM spectrum. In this paper, we explicitly examine the causal and non-causal factors of features within VFMs for DGSS, and propose a simple yet effective method to identify and disentangle them, enabling more robust domain generalization. Specifically, we propose Causal-Tune, a novel fine-tuning strategy designed to extract causal factors and suppress non-causal ones from the features of VFMs. First, we extract the frequency spectrum of features from each layer using the Discrete Cosine Transform (DCT). A Gaussian band-pass filter is then applied to separate the spectrum into causal and non-causal components. To further refine the causal components, we introduce a set of causal-aware learnable tokens that operate in the frequency domain, while the non-causal components are discarded. Finally, refined features are transformed back into the spatial domain via inverse DCT and passed to the next layer. Extensive experiments conducted on various cross-domain tasks demonstrate the effectiveness of Causal-Tune. In particular, our method achieves superior performance under adverse weather conditions, improving +4.8% mIoU over the baseline in snow conditions.
Yin Zhang 0015, Yongqiang Zhang 0007, Yaoyue Zheng, Bogdan Raducanu, Dan Liu 0004
AAAI2
2026 Stream-DINO: exploring DETR-based online object detection with streaming perception
Yongqiang Zhang 0007, Yin Zhang 0015, Zian Zhang, Jinwei Sun
Appl. Intell.2
2026 Parameter-efficient multimodal adaptation for adverse condition depth estimation
Guanglei Yang, Yongqiang Zhang 0007, Zhun Zhong, Wangmeng Zuo
Expert Syst. Appl.3
2026 Plug-and-Play global and local collaborative fusion for weakly supervised object detection
abstract
• We propose a plug-and-play global and local collaborative fusion method to improve the performance of weakly supervised object detection. • We design a pixel-level global information awareness module that utilizes singular value decomposition for image reconstruction. • We propose a local detail fusion module to enable the visual encoder to learn detailed information about target objects. • We demonstrate the effectiveness and superiority of our plug-and-play method through extensive experiments. Weakly supervised object detection (WSOD) has drawn much attention due to its closeness to practical applications, and researchers have proposed the multi-instance learning (MIL) approach to handle it as a multi-class classification problem. Although these methods have yielded promising results, extraneous information in the images severely affects the model’s feature learning due to the lack of instance-level annotation. To alleviate this limitation, in this paper, a global and local collaborative fusion method is proposed for WSOD by leveraging the complementary information of the original image and its low-rank approximation. Specifically, we design a pixel-level global information awareness (GIA) module to reconstruct the input image and remove redundant noise, which are then fed into a visual encoder to extract the features from a global perspective. Moreover, to compensate for the lack of detail preservation in GIA, we further propose a local detail fusion (LDF) module that fuses image details by leveraging both reconstructed and input images. Our proposed GIA-LDF modules are architecture-agnostic and can be seamlessly embedded into any MIL-based WSOD pipeline. Extensive experiments validate the effectiveness of our plug-and-play GIA-LDF for WSOD. We achieve 60.2%, 57.4%, and 23.2% mAP on PASCAL VOC 2007, VOC 2012, and COCO, respectively, surpassing baseline methods by +2.0%, +1.2%, and +0.3%, and establishing new state-of-the-art performance across all benchmarks.
Qiuyu Liang, Yongqiang Zhang 0007
Knowl. Based Syst.2
2026 WPD: Weather prompt driven zero-shot adverse condition depth estimation
Yongqiang Zhang 0007, Zian Zhang, Yin Zhang 0015, Wangmeng Zuo
Pattern Recognit.2
2026 Image signal process with dynamic class-rebalanced and IoU-threshold for unsupervised domain adaptive dark object detection
Yin Zhang 0015, Yongqiang Zhang 0007, Zian Zhang, Mingli Ding, Bogdan Raducanu, Dan Liu 0004
Pattern Recognit.2
2026 ILD: Image-Level Labels Driven Active Learning Object Detection
abstract
Existing SOTA methods in active learning object detection(ALOD) achieve impressive results, but they overlook two problems: (1) the requirement for instance-level labels during initialization, and (2) the constrained localization ability of the pre-trained fully supervised detector in the active learning phase. Problem (1) contradicts the fundamental purpose of active learning in balancing annotation costs and detection performance. Problem (2) arises from the fact that the active learning process relies on a single pre-trained fully supervised detector. To tackle these problems, we propose Image-level Labels Driven active learning object detection (termed as ILD). Specifically, we propose a multi-step reasoning process based on the chain-of-thought only using image-level labels, including a class-number-aware step and an iterative step, to enhance the detection ability of VLM. The detection results of the VLM and weakly supervised detector are used as pseudo ground-truth boxes to initialize a fully supervised detector during AL initialization. Thus, the initialization process of ILD eliminates the requirement for instance-level labels. In the active learning stage, we design two novel uncertainty and diversity acquisition functions to select the most informative images based on collaborative outputs from both the weakly supervised detector and the pre-trained fully supervised detector. The collaborative mechanism jointly measures the uncertainty of two detectors and the diversity of object features, thereby enhancing the localization quality. Extensive experiments demonstrate that the proposed ILD achieves state-of-the-art performance(i.e., 77.5%, 25.7%, and 27.9%) on PASCAL VOC2007, MS COCO2014 and MS COCO2017 datasets, surpassing the SOTA methods by 3.4%, 1.2% and 4.7%, respectively. Our code is publicly available on https://github.com/RuiTianHIT/ILD.
Yongqiang Zhang 0007, Zian Zhang, Yin Zhang 0015, Wangmeng Zuo
IEEE Trans. Circuits Syst. Video Technol.3
2025 SAM based Region-Word Clustering and Inference Score Adjusting for Open-Vocabulary Object Detection
Qiuyu Liang, Yongqiang Zhang 0007
ACM Multimedia2
2025 Revising Representation and Target Deviations for Accurate Human Pose Estimation
abstract
Owing to the normalized instance scales and robust supervision, heatmap-based human pose estimation (HPE) methods with top-down paradigm have achieved a dominant performance. However, there are two inherent deviations in the basic framework, i.e., representation and target deviations, resulting in performance bottlenecks. The representation deviation is caused by transforming various scales of instances into a unified input size, which results in performance degradation because data with different scale-related characteristics can hardly be handled via unified parameters. The target deviation is caused by exploiting a prior distribution (e.g., Gauss) to model the prediction error, which hinders sufficient network training. In this article, we propose a novel framework called DRPose to revise the abovementioned deviations. Specifically, to address the representation deviation, a scale-aware domain bridging (SDB) block is proposed to transfer feature maps from multiple scale-dependent domains into a unified intermediate domain with dynamic parameters. To address the target deviation, a differentiable coordinate decoder (DCD) is presented to adaptively adjust target distribution of heatmaps in an end-to-end manner. Extensive experiments show that the proposed method significantly improves the performance of most existing models with negligible additional cost. Beyond this, our method achieves 77.1% AP on the COCO test-dev set, outperforming prior works with similar model complexity.
Zian Zhang, Yongqiang Zhang 0007, Yancheng Bai, Yin Zhang 0015, Mingli Ding, Wangmeng Zuo
IEEE Trans. Neural Networks Learn. Syst.2
2024 ISP-Teacher: Image Signal Process with Disentanglement Regularization for Unsupervised Domain Adaptive Dark Object Detection
abstract
Object detection in dark conditions has always been a great challenge due to the complex formation process of low-light images. Currently, the mainstream methods usually adopt domain adaptation with Teacher-Student architecture to solve the dark object detection problem, and they imitate the dark conditions by using non-learnable data augmentation strategies on the annotated source daytime images. Note that these methods neglected to model the intrinsic imaging process, i.e. image signal processing (ISP), which is important for camera sensors to generate low-light images. To solve the above problems, in this paper, we propose a novel method named ISP-Teacher for dark object detection by exploring Teacher-Student architecture from a new perspective (i.e. self-supervised learning based ISP degradation). Specifically, we first design a day-to-night transformation module that consistent with the ISP pipeline of the camera sensors (ISP-DTM) to make the augmented images look more in line with the natural low-light images captured by cameras, and the ISP-related parameters are learned in a self-supervised manner. Moreover, to avoid the conflict between the ISP degradation and detection tasks in a shared encoder, we propose a disentanglement regularization (DR) that minimizes the absolute value of cosine similarity to disentangle two tasks and push two gradients vectors as orthogonal as possible. Extensive experiments conducted on two benchmarks show the effectiveness of our method in dark object detection. In particular, ISP-Teacher achieves an improvement of +2.4% AP and +3.3% AP over the SOTA method on BDD100k and SHIFT datasets, respectively. The code can be found at https://github.com/zhangyin1996/ISP-Teacher.
Yin Zhang 0015, Yongqiang Zhang 0007, Zian Zhang, Mingli Ding
AAAI2
2024 R-CCF: region-aware continual contrastive fusion for weakly supervised object detection
Yongqiang Zhang 0007, Yin Zhang 0015, Zian Zhang, Yancheng Bai, Mingli Ding, Wangmeng Zuo
Appl. Intell.1
2024 Towards Non Co-occurrence Incremental Object Detection with Unlabeled In-the-Wild Data
Yongqiang Zhang 0007, Mingli Ding, Gim Hee Lee
Int. J. Comput. Vis.2
2024 Vital information is only worth one thumbnail: Towards efficient human pose estimation
Zian Zhang, Yongqiang Zhang 0007, Yin Zhang 0015, Mingli Ding
Pattern Recognit.2
2023 Incremental-DETR: Incremental Few-Shot Object Detection via Self-Supervised Learning
abstract
Incremental few-shot object detection aims at detecting novel classes without forgetting knowledge of the base classes with only a few labeled training data from the novel classes. Most related prior works are on incremental object detection that rely on the availability of abundant training samples per novel class that substantially limits the scalability to real-world setting where novel data can be scarce. In this paper, we propose the Incremental-DETR that does incremental few-shot object detection via fine-tuning and self-supervised learning on the DETR object detector. To alleviate severe over-fitting with few novel class data, we first fine-tune the class-specific components of DETR with self-supervision from additional object proposals generated using Selective Search as pseudo labels. We further introduce an incremental few-shot fine-tuning strategy with knowledge distillation on the class-specific components of DETR to encourage the network in detecting novel classes without forgetting the base classes. Extensive experiments conducted on standard incremental object detection and incremental few-shot object detection settings show that our approach significantly outperforms state-of-the-art methods by a large margin. Our source code is available at https://github.com/dongnana777/Incremental-DETR.
Yongqiang Zhang 0007, Mingli Ding, Gim Hee Lee
AAAI2
2023 Boosting Long-tailed Object Detection via Step-wise Learning on Smooth-tail Data
abstract
Real-world data tends to follow a long-tailed distribution, where the class imbalance results in dominance of the head classes during training. In this paper, we propose a frustratingly simple but effective step-wise learning framework to gradually enhance the capability of the model in detecting all categories of long-tailed datasets. Specifically, we build smooth-tail data where the long-tailed distribution of categories decays smoothly to correct the bias towards head classes. We pre-train a model on the whole long-tailed data to preserve discriminability between all categories. We then fine-tune the class-agnostic modules of the pre-trained model on the head class dominant replay data to get a head class expert model with improved decision boundaries from all categories. Finally, we train a unified model on the tail class dominant replay data while transferring knowledge from the head class expert model to ensure accurate detection of all categories. Extensive experiments on long-tailed datasets LVIS v0.5 and LVIS v1.0 demonstrate the superior performance of our method, where we can improve the AP with ResNet-50 backbone from 27.0% to 30.3% AP, and especially for the rare categories from 15.5% to 24.9% AP. Our best model using ResNet-101 backbone can achieve 30.7% AP, which suppresses all existing detectors using the same backbone. Our source code is available at https://github.com/dongnana777/Long-tailed-object-detection.
Yongqiang Zhang 0007, Mingli Ding, Gim Hee Lee
ICCV2
2023 Class-incremental object detection
Yongqiang Zhang 0007, Mingli Ding, Yancheng Bai
Pattern Recognit.2
2023 ThumbDet: One thumbnail image is enough for object detection
Yongqiang Zhang 0007, Yin Zhang 0015, Zian Zhang, Yancheng Bai, Wangmeng Zuo, Mingli Ding
Pattern Recognit.1
2023 Uncertainty-Aware Graph-Guided Weakly Supervised Object Detection
abstract
Weakly supervised object detection is an important and challenging task in the computer vision community. In this paper, we treat weakly supervised object detection as a self-training learning task. Based on the framework of self-training, weakly supervised object detection has two uncertainties during training,i.e., the uncertainty of the pseudo labels and the uncertainty of bounding box regression. To this end, we propose an uncertainty-aware graph-guided self-training framework to eliminate these uncertainties. First, we adopt a precise positive and negative sampling strategy to generate pseudo labels to solve the problem of pseudo label uncertainty. Then, we design a weighted location refinement branch based on Bayesian uncertainty modeling to overcome the bounding box regression uncertainty. Moreover, the imbalance between classification and localization tasks prevents the model from generating the task-aware feature map, and redundant proposals, if not handled properly, also introduce uncertainty to the detector. To overcome this problem, we design a graph-guided module that not only balances the two tasks from the perspective of features but also makes full use of proposals. Furthermore, the relation graph of proposals is constructed by clustering proposals, and then, the graph convolution network (GCN) is applied to propagate information on the graph. Thus, accurate feature representations of the objects are obtained through the graph-guided module, and the classification and localization tasks can promote each other. Extensive experiments on the PASCAL VOC 2007 and 2012 datasets demonstrate the effectiveness of our framework, and we obtain 55.2% and 52.0% mAPs on VOC2007 and VOC2012, respectively, showing its superiority over the state-of-the-art approaches by a large margin.
Yueyi Zhu, Yongqiang Zhang 0007, Mingli Ding, Wangmeng Zuo
IEEE Trans. Circuits Syst. Video Technol.2
2022 One-stage object detection knowledge distillation via adversarial learning
Yongqiang Zhang 0007, Mingli Ding, Shibiao Xu, Yancheng Bai
Appl. Intell.2
2022 Bi-directional class-wise adversaries for unsupervised domain adaptation
Guanglei Yang, Mingli Ding, Yongqiang Zhang 0007
Appl. Intell.3
2021 Bridging Non Co-occurrence with Unlabeled In-the-wild Data for Incremental Object Detection
abstract
Deep networks have shown remarkable results in the task of object detection. However, their performance suffers critical drops when they are subsequently trained on novel classes without any sample from the base classes originally used to train the model. This phenomenon is known as catastrophic forgetting. Recently, several incremental learning methods are proposed to mitigate catastrophic forgetting for object detection. Despite the effectiveness, these methods require co-occurrence of the unlabeled base classes in the training data of the novel classes. This requirement is impractical in many real-world settings since the base classes do not necessarily co-occur with the novel classes. In view of this limitation, we consider a more practical setting of complete absence of co-occurrence of the base and novel classes for the object detection task. We propose the use of unlabeled in-the-wild data to bridge the non co-occurrence caused by the missing base classes during the training of additional novel classes. To this end, we introduce a blind sampling strategy based on the responses of the base-class model and pre-trained novel-class model to select a smaller relevant dataset from the large in-the-wild dataset for incremental learning. We then design a dual-teacher distillation framework to transfer the knowledge distilled from the base- and novel-class teacher models to the student model using the sampled in-the-wild data. Experimental results on the PASCAL VOC and MS COCO datasets show that our proposed method significantly outperforms other state-of-the-art class-incremental object detection methods when there is no co-occurrence between the base and novel classes during training.
Yongqiang Zhang 0007, Mingli Ding, Gim Hee Lee
NeurIPS2
2021 KGSNet: Key-Point-Guided Super-Resolution Network for Pedestrian Detection in the Wild
abstract
In real-world scenarios (i.e., in the wild), pedestrians are often far from the camera (i.e., small scale), and they often gather together and occlude with each other (i.e., heavily occluded). However, detecting these small-scale and heavily occluded pedestrians remains a challenging problem for the existing pedestrian detection methods. We argue that these problems arise because of two factors: 1) insufficient resolution of feature maps for handling small-scale pedestrians and 2) lack of an effective strategy for extracting body part information that can directly deal with occlusion. To solve the above-mentioned problems, in this article, we propose a key-point-guided super-resolution network (coined KGSNet) for detecting these small-scale and heavily occluded pedestrians in the wild. Specifically, to address factor 1), a super-resolution network is first trained to generate a clear super-resolution pedestrian image from a small-scale one. In the super-resolution network, we exploit key points of the human body to guide the super-resolution network to recover fine details of the human body region for easier pedestrian detection. To address factor 2), a part estimation module is proposed to encode the semantic information of different human body parts where four semantic body parts (i.e., head and upper/middle/bottom body) are extracted based on the key points. Finally, based on the generated clear super-resolved pedestrian patches padded with the extracted semantic body part images at the image level, a classification network is trained to further distinguish pedestrians/backgrounds from the inputted proposal regions. Both proposed networks (i.e., super-resolution network and classification network) are optimized in an alternating manner and trained in an end-to-end fashion. Extensive experiments on the challenging CityPersons data set demonstrate the effectiveness of the proposed method, which achieves superior performance over previous state-of-the-art methods, especially for those small-scale and heavily occluded instances. Beyond this, we also achieve state-of-the-art performance (i.e., 3.89% MR-2on the reasonable subset) on the Caltech data set.
Yongqiang Zhang 0007, Yancheng Bai, Mingli Ding, Shibiao Xu, Bernard Ghanem
IEEE Trans. Neural Networks Learn. Syst.1
2020 Multi-task Generative Adversarial Network for Detecting Small Objects in the Wild
Yongqiang Zhang 0007, Yancheng Bai, Mingli Ding, Bernard Ghanem
Int. J. Comput. Vis.1
2019 Corrigendum to 'Weakly-supervised Object Detection via Mining Pseudo Ground Truth Bounding-boxes' [Pattern Recognition 84 (2018) 68-81]
Yongqiang Zhang 0007, Yancheng Bai, Mingli Ding, Bernard Ghanem
Pattern Recognit.1
2019 Detecting small faces in the wild based on generative adversarial network and contextual information
Yongqiang Zhang 0007, Mingli Ding, Yancheng Bai, Bernard Ghanem
Pattern Recognit.1
2019 Learning a strong detector for action localization in videos
Yongqiang Zhang 0007, Mingli Ding, Yancheng Bai, Bernard Ghanem
Pattern Recognit. Lett.1
2018 Finding Tiny Faces in the Wild With Generative Adversarial Network
abstract
Face detection techniques have been developed for decades, and one of remaining open challenges is detecting small faces in unconstrained conditions. The reason is that tiny faces are often lacking detailed information and blurring. In this paper, we proposed an algorithm to directly generate a clear high-resolution face from a blurry small one by adopting a generative adversarial network (GAN). Toward this end, the basic GAN formulation achieves it by super-resolving and refining sequentially (e.g. SR-GAN and cycle-GAN). However, we design a novel network to address the problem of super-resolving and refining jointly. We also introduce new training losses to guide the generator network to recover fine details and to promote the discriminator network to distinguish real vs. fake and face vs. non-face simultaneously. Extensive experiments on the challenging dataset WIDER FACE demonstrate the effectiveness of our proposed method in restoring a clear high-resolution face from a blurry small one, and show that the detection performance outperforms other state-of-the-art methods.
Yancheng Bai, Yongqiang Zhang 0007, Mingli Ding, Bernard Ghanem
CVPR2
2018 W2F: A Weakly-Supervised to Fully-Supervised Framework for Object Detection
abstract
Weakly-supervised object detection has attracted much attention lately, since it does not require bounding box annotations for training. Although significant progress has also been made, there is still a large gap in performance between weakly-supervised and fully-supervised object detection. Recently, some works use pseudo ground-truths which are generated by a weakly-supervised detector to train a supervised detector. Such approaches incline to find the most representative parts of objects, and only seek one ground-truth box per class even though many same-class instances exist. To overcome these issues, we propose a weakly-supervised to fully-supervised framework, where a weakly-supervised detector is implemented using multiple instance learning. Then, we propose a pseudo ground-truth excavation (PGE) algorithm to find the pseudo ground-truth of each instance in the image. Moreover, the pseudo ground-truth adaptation (PGA) algorithm is designed to further refine the pseudo ground-truths from PGE. Finally, we use these pseudo ground-truths to train a fully-supervised detector. Extensive experiments on the challenging PASCAL VOC 2007 and 2012 benchmarks strongly demonstrate the effectiveness of our framework. We obtain 52.4% and 47.8% mAP on VOC2007 and VOC2012 respectively, a significant improvement over previous state-of-the-art methods.
Yongqiang Zhang 0007, Yancheng Bai, Mingli Ding, Bernard Ghanem
CVPR1
2018 SOD-MTGAN: Small Object Detection via Multi-Task Generative Adversarial Network
Yancheng Bai, Yongqiang Zhang 0007, Mingli Ding, Bernard Ghanem
ECCV (13)2
2018 Weakly-supervised object detection via mining pseudo ground truth bounding-boxes
Yongqiang Zhang 0007, Yancheng Bai, Mingli Ding, Bernard Ghanem
Pattern Recognit.1
2016 Reading recognition of pointer meter based on pattern recognition and dynamic three-points on a line
abstract
Pointer meters are frequently applied to industrial production for they are directly readable. They should be calibrated regularly to ensure the precision of the readings. Currently the method of manual calibration is most frequently adopted to accomplish the verification of the pointer meter, and professional skills and subjective judgment may lead to big measurement errors and poor reliability and low efficiency, etc. In the past decades, with the development of computer technology, the skills of machine vision and digital image processing have been applied to recognize the reading of the dial instrument. In terms of the existing recognition methods, all the parameters of dial instruments are supposed to be the same, which is not the case in practice. In this work, recognition of pointer meter reading is regarded as an issue of pattern recognition. We obtain the features of a small area around the detected point, make those features as a pattern, divide those certified images based on Gradient Pyramid Algorithm, train a classifier with the support vector machine (SVM) and complete the pattern matching of the divided mages. Then we get the reading of the pointer meter precisely under the theory of dynamic three points make a line (DTPML), which eliminates the error caused by tiny differences of the panels. Eventually, the result of the experiment proves that the proposed method in this work is superior to state-of-the-art works.
Yongqiang Zhang 0007, Mingli Ding, Wuyifang Fu
ICMV1