EDBT 2026 Demo / reviewers in the wild / expert
Mingli Ding
dblp:12/113
· DBLP profile ↗
36ranked-venue papers
1as first author
22since 2021 · last 2026
0000-0002-7510-7043ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 1 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Image signal process with dynamic class-rebalanced and IoU-threshold for unsupervised domain adaptive dark object detection
Yin Zhang 0015, Yongqiang Zhang 0007, Zian Zhang, Mingli Ding, Bogdan Raducanu, Dan Liu 0004 |
Pattern Recognit. | 6 |
| 2025 | Revising Representation and Target Deviations for Accurate Human Pose EstimationabstractOwing to the normalized instance scales and robust supervision, heatmap-based human pose estimation (HPE) methods with top-down paradigm have achieved a dominant performance. However, there are two inherent deviations in the basic framework, i.e., representation and target deviations, resulting in performance bottlenecks. The representation deviation is caused by transforming various scales of instances into a unified input size, which results in performance degradation because data with different scale-related characteristics can hardly be handled via unified parameters. The target deviation is caused by exploiting a prior distribution (e.g., Gauss) to model the prediction error, which hinders sufficient network training. In this article, we propose a novel framework called DRPose to revise the abovementioned deviations. Specifically, to address the representation deviation, a scale-aware domain bridging (SDB) block is proposed to transfer feature maps from multiple scale-dependent domains into a unified intermediate domain with dynamic parameters. To address the target deviation, a differentiable coordinate decoder (DCD) is presented to adaptively adjust target distribution of heatmaps in an end-to-end manner. Extensive experiments show that the proposed method significantly improves the performance of most existing models with negligible additional cost. Beyond this, our method achieves 77.1% AP on the COCO test-dev set, outperforming prior works with similar model complexity. Zian Zhang, Yongqiang Zhang 0007, Yancheng Bai, Yin Zhang 0015, Mingli Ding, Wangmeng Zuo |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | ISP-Teacher: Image Signal Process with Disentanglement Regularization for Unsupervised Domain Adaptive Dark Object DetectionabstractObject detection in dark conditions has always been a great challenge due to the complex formation process of low-light images. Currently, the mainstream methods usually adopt domain adaptation with Teacher-Student architecture to solve the dark object detection problem, and they imitate the dark conditions by using non-learnable data augmentation strategies on the annotated source daytime images. Note that these methods neglected to model the intrinsic imaging process, i.e. image signal processing (ISP), which is important for camera sensors to generate low-light images. To solve the above problems, in this paper, we propose a novel method named ISP-Teacher for dark object detection by exploring Teacher-Student architecture from a new perspective (i.e. self-supervised learning based ISP degradation). Specifically, we first design a day-to-night transformation module that consistent with the ISP pipeline of the camera sensors (ISP-DTM) to make the augmented images look more in line with the natural low-light images captured by cameras, and the ISP-related parameters are learned in a self-supervised manner. Moreover, to avoid the conflict between the ISP degradation and detection tasks in a shared encoder, we propose a disentanglement regularization (DR) that minimizes the absolute value of cosine similarity to disentangle two tasks and push two gradients vectors as orthogonal as possible. Extensive experiments conducted on two benchmarks show the effectiveness of our method in dark object detection. In particular, ISP-Teacher achieves an improvement of +2.4% AP and +3.3% AP over the SOTA method on BDD100k and SHIFT datasets, respectively. The code can be found at https://github.com/zhangyin1996/ISP-Teacher. Yin Zhang 0015, Yongqiang Zhang 0007, Zian Zhang, Mingli Ding |
AAAI | 6 |
| 2024 | R-CCF: region-aware continual contrastive fusion for weakly supervised object detection
Yongqiang Zhang 0007, Yin Zhang 0015, Zian Zhang, Yancheng Bai, Mingli Ding, Wangmeng Zuo |
Appl. Intell. | 6 |
| 2024 | Towards Non Co-occurrence Incremental Object Detection with Unlabeled In-the-Wild Data
Yongqiang Zhang 0007, Mingli Ding, Gim Hee Lee |
Int. J. Comput. Vis. | 3 |
| 2024 | Vital information is only worth one thumbnail: Towards efficient human pose estimation
Zian Zhang, Yongqiang Zhang 0007, Yin Zhang 0015, Mingli Ding |
Pattern Recognit. | 5 |
| 2023 | Incremental-DETR: Incremental Few-Shot Object Detection via Self-Supervised LearningabstractIncremental few-shot object detection aims at detecting novel classes without forgetting knowledge of the base classes with only a few labeled training data from the novel classes. Most related prior works are on incremental object detection that rely on the availability of abundant training samples per novel class that substantially limits the scalability to real-world setting where novel data can be scarce. In this paper, we propose the Incremental-DETR that does incremental few-shot object detection via fine-tuning and self-supervised learning on the DETR object detector. To alleviate severe over-fitting with few novel class data, we first fine-tune the class-specific components of DETR with self-supervision from additional object proposals generated using Selective Search as pseudo labels. We further introduce an incremental few-shot fine-tuning strategy with knowledge distillation on the class-specific components of DETR to encourage the network in detecting novel classes without forgetting the base classes. Extensive experiments conducted on standard incremental object detection and incremental few-shot object detection settings show that our approach significantly outperforms state-of-the-art methods by a large margin. Our source code is available at https://github.com/dongnana777/Incremental-DETR. Yongqiang Zhang 0007, Mingli Ding, Gim Hee Lee |
AAAI | 3 |
| 2023 | Boosting Long-tailed Object Detection via Step-wise Learning on Smooth-tail DataabstractReal-world data tends to follow a long-tailed distribution, where the class imbalance results in dominance of the head classes during training. In this paper, we propose a frustratingly simple but effective step-wise learning framework to gradually enhance the capability of the model in detecting all categories of long-tailed datasets. Specifically, we build smooth-tail data where the long-tailed distribution of categories decays smoothly to correct the bias towards head classes. We pre-train a model on the whole long-tailed data to preserve discriminability between all categories. We then fine-tune the class-agnostic modules of the pre-trained model on the head class dominant replay data to get a head class expert model with improved decision boundaries from all categories. Finally, we train a unified model on the tail class dominant replay data while transferring knowledge from the head class expert model to ensure accurate detection of all categories. Extensive experiments on long-tailed datasets LVIS v0.5 and LVIS v1.0 demonstrate the superior performance of our method, where we can improve the AP with ResNet-50 backbone from 27.0% to 30.3% AP, and especially for the rare categories from 15.5% to 24.9% AP. Our best model using ResNet-101 backbone can achieve 30.7% AP, which suppresses all existing detectors using the same backbone. Our source code is available at https://github.com/dongnana777/Long-tailed-object-detection. Yongqiang Zhang 0007, Mingli Ding, Gim Hee Lee |
ICCV | 3 |
| 2023 | Self-training transformer for source-free domain adaptation
Guanglei Yang, Zhun Zhong, Mingli Ding, Nicu Sebe, Elisa Ricci 0001 |
Appl. Intell. | 3 |
| 2023 | Uncertainty-Aware Contrastive Distillation for Incremental Semantic SegmentationabstractA fundamental and challenging problem in deep learning is catastrophic forgetting, i.e., the tendency of neural networks to fail to preserve the knowledge acquired from old tasks when learning new tasks. This problem has been widely investigated in the research community and several Incremental Learning (IL) approaches have been proposed in the past years. While earlier works in computer vision have mostly focused on image classification and object detection, more recently some IL approaches for semantic segmentation have been introduced. These previous works showed that, despite its simplicity, knowledge distillation can be effectively employed to alleviate catastrophic forgetting. In this paper, we follow this research direction and, inspired by recent literature on contrastive learning, we propose a novel distillation framework, Uncertainty-aware Contrastive Distillation (UCD). In a nutshell, UCDis operated by introducing a novel distillation loss that takes into account all the images in a mini-batch, enforcing similarity between features associated to all the pixels from the same classes, and pulling apart those corresponding to pixels from different classes. In order to mitigate catastrophic forgetting, we contrast features of the new model with features extracted by a frozen model learned at the previous incremental step. Our experimental results demonstrate the advantage of the proposed distillation technique, which can be used in synergy with previous IL approaches, and leads to state-of-art performance on three commonly adopted benchmarks for incremental semantic segmentation. Guanglei Yang, Enrico Fini, Dan Xu 0002, Paolo Rota, Mingli Ding, Moin Nabi, Xavier Alameda-Pineda, Elisa Ricci 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Class-incremental object detection
Yongqiang Zhang 0007, Mingli Ding, Yancheng Bai |
Pattern Recognit. | 3 |
| 2023 | ThumbDet: One thumbnail image is enough for object detection
Yongqiang Zhang 0007, Yin Zhang 0015, Zian Zhang, Yancheng Bai, Wangmeng Zuo, Mingli Ding |
Pattern Recognit. | 7 |
| 2023 | Uncertainty-Aware Graph-Guided Weakly Supervised Object DetectionabstractWeakly supervised object detection is an important and challenging task in the computer vision community. In this paper, we treat weakly supervised object detection as a self-training learning task. Based on the framework of self-training, weakly supervised object detection has two uncertainties during training,i.e., the uncertainty of the pseudo labels and the uncertainty of bounding box regression. To this end, we propose an uncertainty-aware graph-guided self-training framework to eliminate these uncertainties. First, we adopt a precise positive and negative sampling strategy to generate pseudo labels to solve the problem of pseudo label uncertainty. Then, we design a weighted location refinement branch based on Bayesian uncertainty modeling to overcome the bounding box regression uncertainty. Moreover, the imbalance between classification and localization tasks prevents the model from generating the task-aware feature map, and redundant proposals, if not handled properly, also introduce uncertainty to the detector. To overcome this problem, we design a graph-guided module that not only balances the two tasks from the perspective of features but also makes full use of proposals. Furthermore, the relation graph of proposals is constructed by clustering proposals, and then, the graph convolution network (GCN) is applied to propagate information on the graph. Thus, accurate feature representations of the objects are obtained through the graph-guided module, and the classification and localization tasks can promote each other. Extensive experiments on the PASCAL VOC 2007 and 2012 datasets demonstrate the effectiveness of our framework, and we obtain 55.2% and 52.0% mAPs on VOC2007 and VOC2012, respectively, showing its superiority over the state-of-the-art approaches by a large margin. Yueyi Zhu, Yongqiang Zhang 0007, Mingli Ding, Wangmeng Zuo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Continual Attentive Fusion for Incremental Learning in Semantic SegmentationabstractInternational audience Guanglei Yang, Enrico Fini, Dan Xu 0002, Paolo Rota, Mingli Ding, Hao Tang 0005, Xavier Alameda-Pineda, Elisa Ricci 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | An End-to-End Framework for Joint Denoising and Classification of Hyperspectral ImagesabstractImage denoising and classification are typically conducted separately and sequentially according to their respective objectives. In such a setup, where the two tasks are decoupled, the denoising operation does not optimally serve the classification task and sometimes even deteriorates it. We introduce here a unified deep learning framework for joint denoising and classification of high-dimensional images, and we particularly apply it in the framework of hyperspectral imaging. Earlier works on joint image denoising and classification are very scarce, and to the best of our knowledge, no deep learning models were proposed or studied yet for this type of multitask image processing. A key component in our joint learning model is a compound loss function, designed in such a way that the denoising and classification operations benefit each other iteratively during the learning process. Hyperspectral images (HSIs) are particularly challenging for both denoising and classification due to their high dimensionality and varying noise statistics across the bands. We argue that a well-designed end-to-end deep learning framework for joint denoising and classification is superior to current deep learning approaches for processing HSI data, and we substantiate this by results on real HSI images in remote sensing. We experimentally show that the proposed joint learning framework substantially improves the classification performance compared to the common deep learning approaches in HSI processing, and as a by-product, the denoising results are enhanced as well, especially in terms of the semantic content, benefiting from the classification. Xian Li 0001, Mingli Ding, Yanfeng Gu, Aleksandra Pizurica |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | One-stage object detection knowledge distillation via adversarial learning
Yongqiang Zhang 0007, Mingli Ding, Shibiao Xu, Yancheng Bai |
Appl. Intell. | 3 |
| 2022 | Bi-directional class-wise adversaries for unsupervised domain adaptation
Guanglei Yang, Mingli Ding, Yongqiang Zhang 0007 |
Appl. Intell. | 2 |
| 2022 | Fully Group Convolutional Neural Networks for Robust Spectral-Spatial Feature LearningabstractConvolutional neural network (CNN) has been widely applied in hyperspectral image (HSI) classification exhibiting excellent performance. Weak generalization of CNN models to different datasets is a common issue in this domain largely because of limited amount of labeled training samples. In this article, we propose afullygroup convolutional neural network (FGCNN) method that integrates cascades of shuffled group convolutions tailored to different network stages. To our knowledge, this is the first reported full-group CNN model in general, and we design it in particular for robust spectral–spatial classification of HSI. In the primary feature extraction stage, we develop an original multiscale spectral feature extraction approach based on a novel concept of multikernel depthwise convolution that we define in terms of shuffled and importance-weighted group convolution. In the subsequent stage, we introduce a discriminative spectral–spatial feature extraction method with a novel group competition block to capture informative features with relatively few parameters. The final feature fusion stage is defined as a novel lightweight group feature fusion method that sharply reduces fusion weights compared to traditional methods with fully connected layers. Experimental results on three datasets show that the proposed FGCNN yields robust classification accuracy under the same hyperparameter settings compared to the current state-of-the-art. Xian Li 0001, Mingli Ding, Aleksandra Pizurica |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Spectral Feature Fusion Networks With Dual Attention for Hyperspectral Image ClassificationabstractRecent progress in spectral classification is largely attributed to the use of convolutional neural networks (CNNs). While a variety of successful architectures have been proposed, they all extract spectral features from various portions of adjacent spectral bands. In this article, we take a different approach and develop a deep spectral feature fusion method, which extracts both local and interlocal spectral features, capturing thus also the correlations among nonadjacent bands. To our knowledge, this is the first reported deep spectral feature fusion method. Our model is a two-stream architecture, where an intergroup and a groupwise spectral classifier operate in parallel. The interlocal spectral correlation feature extraction is achieved elegantly, by reshaping the input spectral vectors to form the so-called nonadjacent spectral matrices. We introduce the concept of groupwise band convolution to enable the efficient extraction of discriminative local features with multiple kernels adopting the local spectral content. Another important contribution of this work is a novel dual-channel attention mechanism to identify the most informative spectral features. The model is trained in an end-to-end fashion with a joint loss. Experimental results on real datasets demonstrate excellent performance compared with the current state of the art. Xian Li 0001, Mingli Ding, Aleksandra Pizurica |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Transformer-Based Attention Networks for Continuous Pixel-Wise PredictionabstractWhile convolutional neural networks have shown a tremendous impact on various computer vision tasks, they generally demonstrate limitations in explicitly modeling long-range dependencies due to the intrinsic locality of the convolution operation. Initially designed for natural language processing tasks, Transformers have emerged as alternative architectures with innate global self-attention mechanisms to capture long-range dependencies. In this paper, we propose TransDepth, an architecture that benefits from both convolutional neural networks and transformers. To avoid the network losing its ability to capture locallevel details due to the adoption of transformers, we propose a novel decoder that employs attention mechanisms based on gates. Notably, this is the first paper that applies transformers to pixel-wise prediction problems involving continuous labels (i.e., monocular depth prediction and surface normal estimation). Extensive experiments demonstrate that the proposed TransDepth achieves state-of-theart performance on three challenging datasets. Our code is available at: https://github.com/ygjwd12345/TransDepth. Guanglei Yang, Hao Tang 0005, Mingli Ding, Nicu Sebe, Elisa Ricci 0001 |
ICCV | 3 |
| 2021 | Bridging Non Co-occurrence with Unlabeled In-the-wild Data for Incremental Object DetectionabstractDeep networks have shown remarkable results in the task of object detection. However, their performance suffers critical drops when they are subsequently trained on novel classes without any sample from the base classes originally used to train the model. This phenomenon is known as catastrophic forgetting. Recently, several incremental learning methods are proposed to mitigate catastrophic forgetting for object detection. Despite the effectiveness, these methods require co-occurrence of the unlabeled base classes in the training data of the novel classes. This requirement is impractical in many real-world settings since the base classes do not necessarily co-occur with the novel classes. In view of this limitation, we consider a more practical setting of complete absence of co-occurrence of the base and novel classes for the object detection task. We propose the use of unlabeled in-the-wild data to bridge the non co-occurrence caused by the missing base classes during the training of additional novel classes. To this end, we introduce a blind sampling strategy based on the responses of the base-class model and pre-trained novel-class model to select a smaller relevant dataset from the large in-the-wild dataset for incremental learning. We then design a dual-teacher distillation framework to transfer the knowledge distilled from the base- and novel-class teacher models to the student model using the sampled in-the-wild data. Experimental results on the PASCAL VOC and MS COCO datasets show that our proposed method significantly outperforms other state-of-the-art class-incremental object detection methods when there is no co-occurrence between the base and novel classes during training. Yongqiang Zhang 0007, Mingli Ding, Gim Hee Lee |
NeurIPS | 3 |
| 2021 | KGSNet: Key-Point-Guided Super-Resolution Network for Pedestrian Detection in the WildabstractIn real-world scenarios (i.e., in the wild), pedestrians are often far from the camera (i.e., small scale), and they often gather together and occlude with each other (i.e., heavily occluded). However, detecting these small-scale and heavily occluded pedestrians remains a challenging problem for the existing pedestrian detection methods. We argue that these problems arise because of two factors: 1) insufficient resolution of feature maps for handling small-scale pedestrians and 2) lack of an effective strategy for extracting body part information that can directly deal with occlusion. To solve the above-mentioned problems, in this article, we propose a key-point-guided super-resolution network (coined KGSNet) for detecting these small-scale and heavily occluded pedestrians in the wild. Specifically, to address factor 1), a super-resolution network is first trained to generate a clear super-resolution pedestrian image from a small-scale one. In the super-resolution network, we exploit key points of the human body to guide the super-resolution network to recover fine details of the human body region for easier pedestrian detection. To address factor 2), a part estimation module is proposed to encode the semantic information of different human body parts where four semantic body parts (i.e., head and upper/middle/bottom body) are extracted based on the key points. Finally, based on the generated clear super-resolved pedestrian patches padded with the extracted semantic body part images at the image level, a classification network is trained to further distinguish pedestrians/backgrounds from the inputted proposal regions. Both proposed networks (i.e., super-resolution network and classification network) are optimized in an alternating manner and trained in an end-to-end fashion. Extensive experiments on the challenging CityPersons data set demonstrate the effectiveness of the proposed method, which achieves superior performance over previous state-of-the-art methods, especially for those small-scale and heavily occluded instances. Beyond this, we also achieve state-of-the-art performance (i.e., 3.89% MR-2on the reasonable subset) on the Caltech data set. Yongqiang Zhang 0007, Yancheng Bai, Mingli Ding, Shibiao Xu, Bernard Ghanem |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Bi-Directional Generation for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation facilitates the unlabeled target domain relying on well-established source domain information. The conventional methods forcefully reducing the domain discrepancy in the latent space will result in the destruction of intrinsic data structure. To balance the mitigation of domain gap and the preservation of the inherent structure, we propose a Bi-Directional Generation domain adaptation model with consistent classifiers interpolating two intermediate domains to bridge source and target domains. Specifically, two cross-domain generators are employed to synthesize one domain conditioned on the other. The performance of our proposed method can be further enhanced by the consistent classifiers and the cross-domain alignment constraints. We also design two classifiers which are jointly optimized to maximize the consistency on target sample prediction. Extensive experiments verify that our proposed model outperforms the state-of-the-art on standard cross domain visual benchmarks. Guanglei Yang, Haifeng Xia, Mingli Ding, Zhengming Ding |
AAAI | 3 |
| 2020 | Multi-task Generative Adversarial Network for Detecting Small Objects in the Wild
Yongqiang Zhang 0007, Yancheng Bai, Mingli Ding, Bernard Ghanem |
Int. J. Comput. Vis. | 3 |
| 2020 | Deep Feature Fusion via Two-Stream Convolutional Neural Network for Hyperspectral Image ClassificationabstractThe representation power of convolutional neural network (CNN) models for hyperspectral image (HSI) analysis is in practice limited by the available amount of the labeled samples, which is often insufficient to sustain deep networks with many parameters. We propose a novel approach to boost the network representation power with a two-stream 2-D CNN architecture. The proposed method extracts simultaneously, the spectral features and local spatial and global spatial features, with two 2-D CNN networks and makes use of channel correlations to identify the most informative features. Moreover, we propose a layer-specific regularization and a smooth normalization fusion scheme to adaptively learn the fusion weights for the spectral-spatial features from the two parallel streams. An important asset of our model is the simultaneous training of the feature extraction, fusion, and classification processes with the same cost function. Experimental results on several hyperspectral data sets demonstrate the efficacy of the proposed method compared with the state-of-the-art methods in the field. Xian Li 0001, Mingli Ding, Aleksandra Pizurica |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Group Convolutional Neural Networks for Hyperspectral Image ClassificationabstractConvolutional Neural Network (CNN) has been widely applied in hyperspectral image (HSI) classification exhibiting excellent performance. The CNN model overfitting is a common issue in this domain due to limited amount of labelled training samples. In addition, making the full use of spectral information is still considered an open problem. In this paper, we propose a novel group 2D-CNN model for spectral-spatial classification. Specifically, we propose an original multi-scale spectral feature extraction approach based on a novel concept of multi-kernel depthwise convolution. Furthermore, we exploit for the first time shuffle operation on the group convolutions in HSI spectral-spatial feature extraction to effectively limit the amount of learning parameters. As a result, we design a small and efficient network for HSI classification. Experimental results on real data demonstrate favourable performance compared to the current state-of-the-art. Xian Li 0001, Mingli Ding, Aleksandra Pizurica |
ICIP | 2 |
| 2019 | Corrigendum to 'Weakly-supervised Object Detection via Mining Pseudo Ground Truth Bounding-boxes' [Pattern Recognition 84 (2018) 68-81]
Yongqiang Zhang 0007, Yancheng Bai, Mingli Ding, Bernard Ghanem |
Pattern Recognit. | 3 |
| 2019 | Detecting small faces in the wild based on generative adversarial network and contextual information
Yongqiang Zhang 0007, Mingli Ding, Yancheng Bai, Bernard Ghanem |
Pattern Recognit. | 2 |
| 2019 | Learning a strong detector for action localization in videos
Yongqiang Zhang 0007, Mingli Ding, Yancheng Bai, Bernard Ghanem |
Pattern Recognit. Lett. | 2 |
| 2018 | Finding Tiny Faces in the Wild With Generative Adversarial NetworkabstractFace detection techniques have been developed for decades, and one of remaining open challenges is detecting small faces in unconstrained conditions. The reason is that tiny faces are often lacking detailed information and blurring. In this paper, we proposed an algorithm to directly generate a clear high-resolution face from a blurry small one by adopting a generative adversarial network (GAN). Toward this end, the basic GAN formulation achieves it by super-resolving and refining sequentially (e.g. SR-GAN and cycle-GAN). However, we design a novel network to address the problem of super-resolving and refining jointly. We also introduce new training losses to guide the generator network to recover fine details and to promote the discriminator network to distinguish real vs. fake and face vs. non-face simultaneously. Extensive experiments on the challenging dataset WIDER FACE demonstrate the effectiveness of our proposed method in restoring a clear high-resolution face from a blurry small one, and show that the detection performance outperforms other state-of-the-art methods. Yancheng Bai, Yongqiang Zhang 0007, Mingli Ding, Bernard Ghanem |
CVPR | 3 |
| 2018 | W2F: A Weakly-Supervised to Fully-Supervised Framework for Object DetectionabstractWeakly-supervised object detection has attracted much attention lately, since it does not require bounding box annotations for training. Although significant progress has also been made, there is still a large gap in performance between weakly-supervised and fully-supervised object detection. Recently, some works use pseudo ground-truths which are generated by a weakly-supervised detector to train a supervised detector. Such approaches incline to find the most representative parts of objects, and only seek one ground-truth box per class even though many same-class instances exist. To overcome these issues, we propose a weakly-supervised to fully-supervised framework, where a weakly-supervised detector is implemented using multiple instance learning. Then, we propose a pseudo ground-truth excavation (PGE) algorithm to find the pseudo ground-truth of each instance in the image. Moreover, the pseudo ground-truth adaptation (PGA) algorithm is designed to further refine the pseudo ground-truths from PGE. Finally, we use these pseudo ground-truths to train a fully-supervised detector. Extensive experiments on the challenging PASCAL VOC 2007 and 2012 benchmarks strongly demonstrate the effectiveness of our framework. We obtain 52.4% and 47.8% mAP on VOC2007 and VOC2012 respectively, a significant improvement over previous state-of-the-art methods. Yongqiang Zhang 0007, Yancheng Bai, Mingli Ding, Bernard Ghanem |
CVPR | 3 |
| 2018 | SOD-MTGAN: Small Object Detection via Multi-Task Generative Adversarial Network
Yancheng Bai, Yongqiang Zhang 0007, Mingli Ding, Bernard Ghanem |
ECCV (13) | 3 |
| 2018 | Weakly-supervised object detection via mining pseudo ground truth bounding-boxes
Yongqiang Zhang 0007, Yancheng Bai, Mingli Ding, Bernard Ghanem |
Pattern Recognit. | 3 |
| 2016 | Reading recognition of pointer meter based on pattern recognition and dynamic three-points on a lineabstractPointer meters are frequently applied to industrial production for they are directly readable. They should be calibrated regularly to ensure the precision of the readings. Currently the method of manual calibration is most frequently adopted to accomplish the verification of the pointer meter, and professional skills and subjective judgment may lead to big measurement errors and poor reliability and low efficiency, etc. In the past decades, with the development of computer technology, the skills of machine vision and digital image processing have been applied to recognize the reading of the dial instrument. In terms of the existing recognition methods, all the parameters of dial instruments are supposed to be the same, which is not the case in practice. In this work, recognition of pointer meter reading is regarded as an issue of pattern recognition. We obtain the features of a small area around the detected point, make those features as a pattern, divide those certified images based on Gradient Pyramid Algorithm, train a classifier with the support vector machine (SVM) and complete the pattern matching of the divided mages. Then we get the reading of the pointer meter precisely under the theory of dynamic three points make a line (DTPML), which eliminates the error caused by tiny differences of the panels. Eventually, the result of the experiment proves that the proposed method in this work is superior to state-of-the-art works. Yongqiang Zhang 0007, Mingli Ding, Wuyifang Fu |
ICMV | 2 |
| 2007 | A New Method for Accelerometer Dynamic Compensation Based on CMAC
Mingli Ding, Qingdong Zhou |
ISNN (1) | 1 |
| 2007 | Blind Source Separation in Post-nonlinear Mixtures Using Natural Gradient Descent and Particle Swarm Optimization Algorithm
Kai Song 0001, Mingli Ding, Qi Wang 0034 |
ISNN (3) | 2 |