VLDB 2026 Research / reviewers in the wild / expert
Yongbin Zheng
dblp:96/8575
· DBLP profile ↗
16ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object DetectionabstractTo identify objects beyond predefined categories, open-vocabulary aerial object detection (OVAD) leverages the zero-shot capabilities of visual-language models (VLMs) to generalize from base to novel categories. Existing approaches typically utilize self-learning mechanisms with weak text supervision to generate region-level pseudo-labels to align detectors with VLMs semantic spaces. However, text dependence induces semantic bias, restricting open-vocabulary expansion to text-specified concepts. We propose VK-Det, a visual knowledge-guided open-vocabulary object detection framework without extra supervision. First, we discover and leverage vision encoder's inherent informative region perception to attain fine-grained localization and adaptive distillation. Second, we introduce a novel prototype-aware pseudo-labeling strategy. It models inter-class decision boundaries through feature clustering and maps detection regions to latent categories via prototype matching. This enhances attention to novel objects while compensating for missing supervision. Extensive experiments show state-of-the-art performance, achieving 30.1 mAPᴺ on DIOR and 23.3 mAPᴺ on DOTA, outperforming even extra supervised methods. Jianhang Yao, Yongbin Zheng, Siqi Lu, Wanying Xu |
AAAI | 2 |
| 2026 | Completing Missing Entities: Exploring Consistency Reasoning for Remote Sensing Object DetectionabstractRecent studies in remote sensing object detection have made excellent progress and shown promising performance. However, most current detectors only explore rotation-invariant feature extraction but disregard the valuable spatial and semantic prior knowledge in remote sensing images (RSIs), which limits the detection performance when encountering blurred or heavy occluded objects. To address this issue, we propose a mask-reconstruction relation learning (MRRL) framework to learn such prior knowledge among objects and a consistency-reasoning transformer over relation proposals (CTRP) to recognize objects with limited visual features via consistency reasoning. Specifically, MRRL framework applies random mask to some objects in the training dataset and performs masked objects reconstruction to guide the network to learn the distribution consistency of objects. CTRP is the core component of the MRRL framework, which models the interaction between spatial and semantic priors, and uses easy detected objects to reason hard detected objects. The trained CTRP can be integrated into the existing detector to improve the ability of object detection with limited visual features in RSIs. Extensive experiments on widely-used datasets for two distinct tasks, namely remote sensing object detection task and occluded object detection task, demonstrate the effectiveness of the proposed method. Source code is available at https://github.com/sunpeng96/CTRP_mmrotate. Yongbin Zheng, Wanying Xu, Jian Li 0003, Jiansong Yang |
IEEE Trans. Image Process. | 2 |
| 2025 | Lifting the Structural Morphing for Wide-Angle Images Rectification: Unified Content and Boundary Modeling
Wenting Luan, Siqi Lu, Yongbin Zheng, Wanying Xu, Lang Nie, Zongtan Zhou, Kang Liao |
ICCV | 3 |
| 2024 | Weakly-Supervised Depth Completion during Robotic Micromanipulation from a Monocular Microscopic ImageabstractObtaining three-dimensional information, especially the z-axis depth information, is crucial for robotic micromanipulation. Due to the unavailability of depth sensors such as lidars in micromanipulation setups, traditional depth acquisition methods such as depth from focus or depth from defocus directly infer depth from microscopic images and suffer from poor resolution. Alternatively, micromanipulation tasks obtain accurate depth information by detecting the contact between an end-effector and an object (e.g., a cell). Despite its high accuracy, only sparse depth data can be obtained due to its low efficiency. This paper aims to address the challenge of acquiring dense depth information during robotic cell micromanipulation. A weakly-supervised depth completion network is proposed to take cell images and sparse depth data obtained by contact detection as input to generate a dense depth map. A two-stage data augmentation method is proposed to augment the sparse depth data, and the depth map is optimized by a network refinement method. The experimental results show that the MAE value of the depth prediction error is less than 0.3 µm, which proves the accuracy and effectiveness of the method. This deep learning network pipeline can be seamlessly integrated with the robotic micromanipulation tasks to provide accurate depth information. Yufei Jin, Guanqiao Shan, Yongbin Zheng, Jiangfan Yu, Yu Sun 0001, Zhuoran Zhang 0001 |
ICRA | 5 |
| 2024 | Efficient-PIP: Large-scale Pixel-level Aligned Image Pair Generation for Cross-time Infrared-RGB TranslationabstractGenerative models are gaining momentum in both academic and industrial applications driven by the availability of large-scale datasets, especially in tasks involving Image-to-Image Translation. Meanwhile, poor human perception of nighttime environment has led to a demand for translation from night-vision infrared to day-vision RGB images. However, collecting such cross-modal training data at the same time is impossible due to the thermal imaging properties of infrared cameras, the challenge lies in constructing image pairs during the day and at night respectively, where the requirement for data alignment poses significant difficulties. In this paper, we propose a Pixel-level aligned Image Pair generation framework PIP to explore efficient colorization of high-resolution infrared images. Specifically, we first construct a 3D high-precision point cloud map for the purpose of establishing the correlation between day and night scenes. Corresponding point clouds of modal images are collected simultaneously during data acquisition to obtain image sensor poses by Global Matching with the map, which allows us to calculate the transformation relationship from infrared to RGB image coordinate systems based on the sensor parameters and depth information of the map. Leveraging the relationship, the pixel values of RGB image is projected onto the infrared image followed by optimization as the colored image. Accordingly, we present a dataset NUDT-PIP, the first of its kind containing large-scale pixel-level aligned cross-time infrared-RGB image pairs of complicated real road scenes. Experimental results demonstrate the reliability and strong applicability of our dataset in Image-to-Image Translation. Our code will be released at https://github.com/wjjjjyourFA/NUDT-PIP. Jian Li 0003, Kexin Fei, Bokai Liu, Zongtan Zhou, Yongbin Zheng, Zhenping Sun |
IROS | 7 |
| 2024 | A novel and efficient model pruning method for deep convolutional neural networks by evaluating the direct and indirect effects of filters
Yongbin Zheng, Wanying Xu |
Neurocomputing | 1 |
| 2024 | Target Before Shooting: Accurate Anomaly Detection and Localization Under One Millisecond via Cascade Patch RetrievalabstractIn this work, by re-examining the "matching" nature of Anomaly Detection (AD), we propose a novel AD framework that simultaneously enjoys new records of AD accuracy and dramatically high running speed. In this framework, the anomaly detection problem is solved via a cascade patch retrieval procedure that retrieves the nearest neighbors for each test image patch in a coarse-to-fine fashion. Given a test sample, the top-K most similar training images are first selected based on a robust histogram matching process. Secondly, the nearest neighbor of each test patch is retrieved over the similar geometrical locations on those "most similar images", by using a carefully trained local metric. Finally, the anomaly score of each test image patch is calculated based on the distance to its "nearest neighbor" and the "non-background" probability. The proposed method is termed "Cascade Patch Retrieval" (CPR) in this work. Different from the previous patch-matching-based AD algorithms, CPR selects proper "targets" (reference images and patches) before "shooting" (patch-matching). On the well-acknowledged MVTec AD, BTAD and MVTec-3D AD datasets, the proposed algorithm consistently outperforms all the comparing SOTA methods by remarkable margins, measured by various AD metrics. Furthermore, CPR is extremely efficient. It runs at the speed of 113 FPS with the standard setting while its simplified version only requires less than 1 ms to process an image at the cost of a trivial accuracy drop. The code of CPR is available at https://github.com/flyinghu123/CPR. Jianfei Hu, Bo Li 0090, Hao Chen 0041, Yongbin Zheng, Chunhua Shen |
IEEE Trans. Image Process. | 5 |
| 2022 | A Robust and Accurate End-to-End Template Matching Method Based on the Siamese NetworkabstractTemplate matching is an important and challenging task in remote sensing and computer vision. Existing template matching methods often fail in the presence of complex nonrigid deformation, occlusion, and background clutter. In this letter, inspired by Siamese trackers, we propose an end-to-end template matching method that is based on the Siamese network. Different from the traditional template matching methods, our method treats the template matching task as a classification-regression task. It is more robust to background clutter, occlusion, and nonrigid deformation. Moreover, we introduce a channel-attention mechanism in the cross correlation operation and replace the commonly used intersection-over-union (IoU) with distance-IoU (DIoU) to build a new regression loss, which further improves the performance of our method. Extensive experiments on the commonly used public benchmark demonstrate that our method achieves the state-of-the-art performance. Yongbin Zheng, Wanying Xu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | Image stitching method by multi-feature constrained alignment and colour adjustmentabstractAbstract Image alignment and colour consistency are two challenging tasks for image stitching. Traditional point correspondence methods are difficult to achieve good alignments due to their insufficiency and unreliability. The results are prone to errors and distortions. On the other hand, the problem of colour inconsistency in overlapping area between image pairs is still difficult to solve, especially when the illumination difference between images is large. To solve these problems, the authors integrate point features and line features into a warping model through a designed energy function. Line features will provide geometric constraints for image stitching, and remedy the defect of point correspondences in low‐textured image stitching. A global colour consistency optimization method with colour mapping via a histogram extreme point‐matching algorithm is proposed. The colour characteristic of reference images will be transferred to the others to achieve a global colour consistency. The proposed method is evaluated on a series of images, and compared with other methods. The experiments demonstrate that the proposed method provides convincing stitching results and achieves satisfied colour consistency results. Xingsheng Yuan, Yongbin Zheng, Jiongming Su, Jianzhai Wu |
IET Image Process. | 2 |
| 2020 | R4 Det: Refined single-stage detector with feature recursion and refinement for rotating object detection in aerial images
Yongbin Zheng, Zongtan Zhou, Wanying Xu |
Image Vis. Comput. | 2 |
| 2017 | Visual Tracking via Probabilistic Hypergraph RankingabstractOnline object tracking is a challenging issue because the appearance of an object tends to change due to intrinsic or extrinsic factors. In this paper, we propose a tracking algorithm based on probabilistic hypergraph ranking. First, three types of hypergraphs are constructed to encode local affinity information. Then, a probabilistic hypergraph is built by combining three distinct hypergraphs linearly. Second, an adaptive template constraint is proposed to effectively use the discriminative information of different templates. Third, object tracking is formulated as a transductive learning issue, and the optimal target location is determined by maximum a posteriori estimation on the ranking scores. Finally, a dynamic updating scheme of positive and negative template sets provides the proposed tracker with robustness against appearance variations. A series of experiments and evaluations on various challenging image sequences is performed, and the results show that the proposed algorithm performs favorably against other state-of-the-art methods. Ruitao Lu, Wanying Xu, Yongbin Zheng, Xinsheng Huang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2011 | Robust Visual Tracking by Integrating Lucas-Kanade into Mean-ShiftabstractThe mean-shift algorithm has achieved considerable success in object tracking due to its simplicity and robustness. However, the lack of template update often leads to out of adaptation to affine transformation of the object. The Lucas-Kanade algorithm has some advantages in obtaining the affine parameters. In this paper, we introduce the inverse compositional algorithm, which is equivalent to but more efficient than Lucas-Kanade algorithm, to complement the traditional mean-shift algorithm. In this method, the average of squared error (ASE) between the initial template and the object image which is warped through the obtained affine parameters is computed to decide whether to update the current template. Experimental results show that the mean-shift tracking with Lucas-Kanade algorithm (MSLK) has high tracking accuracy and good robustness to the change of appearance of the object. Lurong Shen, Xinsheng Huang, Wanying Xu, Yongbin Zheng |
ICIG | 4 |
| 2011 | Local Dominant Orientation Based Mutual Information for Multisensor Template MatchingabstractMutual information (MI) has been very successful in multisensor or multimodal image matching. However, it may lead to mismatching due to lack of spacial information. In this paper, based on a local dominant orientation (LDO), which is a stable nature among images of different sensors and is widely used in the relative rotation estimation, an improved MI for multisensor images matching is proposed. Firstly, the frequently used intensity images are converted to a LDO represented form, where the LDO for each pixel is calculated by cumulating the surrounding gradient vectors within a disk like region. Next, we introduce a simple clustering to cluster each transformed image, thus the joint histogram of MI in the matching stage can be reduced significantly, and hence the computations, memory consumption. Our approach is evaluated by 10 groups of multisensor images, and the results have demonstrated its outstanding performances. Yuzhuang Yan, Yongbin Zheng, Wanying Xu, Xinsheng Huang |
ICIG | 2 |
| 2010 | Pyramid Center-Symmetric Local Binary/Trinary Patterns for Effective Pedestrian Detection
Yongbin Zheng, Chunhua Shen, Richard I. Hartley, Xinsheng Huang |
ACCV (4) | 1 |
| 2010 | Pedestrian Detection Using Center-Symmetric Local Binary Patterns
Yongbin Zheng, Chunhua Shen, Xinsheng Huang |
ICIP | 1 |
| 2009 | A New Method for Motion-Blurred Image Blind Restoration Based on Huber Markov Random FieldabstractIn this paper, we are interested in the problem of motion-blurred image blind restoration. A new method for this ill-posed problem is proposed. We present an adaptive Huber Markov Random Field (HMRF) image prior model as the regularization term, which can be suitable for motion-blurred situation, then turn the ill-posed problem to well-posed. It can preserve fine image details and edges. However, image processing always represents as high-dimension equations that are complicated and computationally expensive for stable solutions. To this point, we propose a combinatorial optimization method benefited from variable substitution optimization technique and Tikhonov regularization technique. These two methods implement alternately in frequency domain, hence optimal original image and point spread function (PSF) can be obtained respectively. Some experiments are presented by comparing the proposed method among classical Wiener filter, Richardson-Lucy deconvolution and a state-of-the art method. These experiments demonstrate the advantages of the proposed priori and combinatorial optimization method. Xinsheng Huang, Yuzhuang Yan, Yongbin Zheng |
ICIG | 4 |