EDBT 2026 Demo / reviewers in the wild / expert
Jianzhong Cao
dblp:22/10764
· DBLP profile ↗
20ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0002-1853-3396ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 8 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Frequency-enhanced representation and cost aggregation for multi-view stereo
Jianzhong Cao, Gaopeng Zhang, Boxue Zhang, Weining Chen |
Signal Process. | 2 |
| 2025 | CR-YOLOv8-Based Detection Method for Identifying Non-Functional Satellite ComponentsabstractABSTRACT Detecting non‐functional satellite components is critical for on‐orbit servicing. Current detection methods struggle with complex image noise, motion blur in space environments, and the limited realism of artificially synthesised sample data. To address these challenges, we propose an enhanced you only look once version 8 (YOLOv8)‐based method. In terms of network architecture, we introduce innovative designs for the backbone and neck components. A novel hybrid attention mechanism replaces the conventional approach, improving the perception and processing of intricate image features and significantly enhancing feature extraction. Additionally, we integrate modules inspired by residual networks into the neck structure, improving training adaptability and ensuring robust information transmission. This design highlights key target features while minimising feature attenuation. We also establish the satellite key element (SAKE) dataset under simulated real space conditions, including image noise and jitter blur. This dataset features components such as satellite bodies and solar panels and uses an encoder–decoder network architecture to refine context information. By merging this with a branch preserving high‐resolution details, we enhance dataset expressiveness. Experiments demonstrate that the enhanced algorithm achieves a mean average precision (mAP) of 78.98% on the SAKE dataset, a 2.57% improvement over the original YOLOv8. The refined model effectively detects critical satellite components, showing superior performance in noisy and blurry scenarios. He Bian, Derui Zhang, Jianzhong Cao, Gaopeng Zhang |
IET Image Process. | 6 |
| 2024 | An efficient multi-scale transformer for satellite image dehazingabstractAbstract Given the impressive achievement of convolutional neural networks (CNNs) in grasping image priors from extensive datasets, they have been widely utilized for tasks related to image restoration. Recently, there is been significant progress in another category of neural architectures—Transformers. These models have demonstrated remarkable performance in natural language tasks and higher‐level vision applications. Despite their ability to address some of CNNs limitations, such as restricted receptive fields and adaptability issues, Transformer models often face difficulties when processing images with a high level of detail. This is because the complexity of the computations required increases significantly with the image's spatial resolution. As a result, their application to most high‐resolution image restoration tasks becomes impractical. In our research, we introduce a novel Transformer model, named DehFormer, by implementing specific design modifications in its fundamental components, for example, the multi‐head attention and feed‐forward network. Specifically, the proposed architecture consists of the three modules, that is, (a) multi‐scale feature aggregation network (MSFAN), (b) the gated‐Dconv feed‐forward network (GFFN), (c) and the multi‐Dconv head transposed attention (MDHTA). For the MDHTA module, our objective is to scrutinize the mechanics of scaled dot‐product attention through the utilization of per‐element product operations, thereby bypassing the need for matrix multiplications and operating directly in the frequency domain for enhanced efficiency. For the GFFN module, which enables only the relevant and valuable information to advance through the network hierarchy, thereby enhancing the efficiency of information flow within the model. Extensive experiments are conducted on the SateHazelk, RS‐Haze, and RSID datasets, resulting in performance that significantly exceeds that of existing methods. Jianzhong Cao, Weining Chen |
Expert Syst. J. Knowl. Eng. | 2 |
| 2023 | Feature spatial pyramid network for low-light image enhancement
Xijuan Song, Jijiang Huang, Jianzhong Cao |
Vis. Comput. | 3 |
| 2023 | IPCS: An improved corner detector with intensity, pattern, curvature, and scale
Changlin Wan, Jianzhong Cao, Jingqiu Huang, Deming Xu |
Vis. Comput. | 2 |
| 2022 | Adaptive feedback connection with a single-level feature for object detectionabstractAbstract From the perspective of detector optimisation, detecting objects using only a one‐level feature cannot provide good performance for a wide range of scales. Various complex feature pyramidal structures address this problem using the divide‐and‐conquer strategy and multi‐scale feature fusion. However, this requires adding too many additional convolutional layers and fusion operations. To address the issue, a simple detection part is proposed, which includes three components, namely a one‐level feature map for detection, the encoder structure with feedback connection, and a decoupled head. The redesigned encoder and decoupled head can successfully address the performance decline caused by the one‐level feature‐based detection. Moreover, the proposed method can accelerate the convergence of the detector and achieve a faster inference time. Based on the optimised detection part, an adaptive feedback connection with a single‐level feature (AFS) is proposed for object detection. The experiments conducted on the MS COCO 2017 benchmark show that the proposed method can achieve comparable results with its multi‐scale pyramid counterpart, You Only Look Once v4 (YOLOv4). In addition, AFS can help the YOLOv4 achieve 44.9 mAP at 27 frame per second and converging 82 epochs earlier under the image size of 608×608, which represents a 42.1% improvements in the convergence speed. Zhongling Ruan, Jianzhong Cao, Huinan Guo, Xin Yang 0011 |
IET Comput. Vis. | 2 |
| 2022 | Cross-scale feature fusion connection for a YOLO detectorabstractAbstract Multi‐scale feature fusion is often used to address the issue of scale variations in object detection. However, most of the proposed network architectures only combine the features of two adjacent levels sequentially, so the first fusion nodes in both top‐down and bottom‐up pathways must be blank nodes that only have one input with no feature fusion. In this work, cross‐scale feature fusion connection (CFFC) is proposed which aims to enhance the entire feature hierarchy by propagating the features of each level more efficiently. The proposed method reuses and aggregates all the features of other scales to the blank nodes in both top‐down and bottom‐up pathways. Furthermore, the authors remove the 1 × 1 convolutional layer and replace the shortcut with concatenation before fusing multiple features. These concatenated feature maps are then supervised by the channel attention block at the fusion nodes. This modification allows the network to learn the important degree of each level in concatenated feature maps along the channel dimension. It is also observed that the proposed method alleviates the inconsistency in feature pyramids with fewer parameters. The performance of a YOLO object detector equipped with the proposed method on the COCO test‐dev 2017 is evaluated. The results show that the proposed method outperforms other architectures presented in the literature. Zhongling Ruan, Jianzhong Cao, Hongbo Zhang 0002 |
IET Comput. Vis. | 3 |
| 2022 | Rotation-aware correlation filters for robust visual tracking
Jiawen Liao, Chun Qi, Jianzhong Cao, Long Ren, Chaoning Zhang |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | Infrared and visible image fusion based on edge-preserving guided filter and infrared feature decomposition
Long Ren, Zhibin Pan, Jianzhong Cao |
Signal Process. | 3 |
| 2021 | Temporal Constraint Background-Aware Correlation Filter With Saliency MapabstractCorrelation filter (CF) based trackers have recently drawn great attention in visual tracking community due to their impressive performance and computational efficiency on benchmark datasets. However, the performance of most existing trackers using correlation filter is hampered by two aspects: i) Included background information in the selected rectangular target patch is considered as part of the target, and they are treated as important as the real target in training new filter model, it causes the filter easily drift when target shape changes dramatically. ii) Existing filters use a moving average operation with an empirical weight to update the filter model in each frame, such per frame adaptation constantly introduces new information of the target patch, but never consider the consistence of the historical information and the newly obtained one, further increases the risk of drifting. This paper presents a new framework including saliency map and a novel CF regression model. We reformulate the original optimization problem, and provide a closed form solution for multidimensional features which is solved efficiently using alternating direction method of multipliers (ADMM) and accelerated using Sherman-Morrison lemma, our algorithm as a new framework can be easily integrated into CF base trackers to boost their tracking performance. We perform comprehensive experiments on five benchmarks: OTB-2015, VOT2016, VOT2018, UAV123, and TempleColor-128. Results show that the proposed method performs favorably against lots of state-of-the-art methods with a speed close to real-time. Our method with deep features performs much better on all 5 datasets. Our code will be released to facilitate further researches. Jiawen Liao, Chun Qi, Jianzhong Cao |
IEEE Trans. Multim. | 3 |
| 2021 | A contour self-compensated network for salient object detection
Jianzhong Cao |
Vis. Comput. | 3 |
| 2020 | Visual Tracking Via Temporally-Regularized Context-Aware Correlation FiltersabstractClassical discriminative correlation filter (DCF) model suffers from boundary effects, several modified discriminative correlation filter models have been proposed to mitigate this drawback using enlarged search region, and remarkable performance improvement has been reported by related papers. However, model deterioration is still not well addressed when facing occlusion and other challenging scenarios. In this work, we propose a novel Temporally-regularized Context-aware Correlation Filters (TCCF) model to model the target appearance more robustly. We take advantage of the enlarged search region to obtain more negative samples to make the filter sufficiently trained, and a temporal regularizer, which restricting variation in filter models between frames, is seamlessly integrated into the original formulation. Our model is derived from the new discriminative learning loss formulation, a closed form solution for multidimensional features is provided, which is solved efficiently using Alternating Direction Method of Multipliers (ADMM). Extensive experiments on standard OTB-2015, TempleColor-128 and VOT-2016 benchmarks show that the proposed approach performs favorably against many state-of-the-art methods with real-time performance of 28fps on single CPU. Jiawen Liao, Chun Qi, Jianzhong Cao, He Bian |
ICIP | 3 |
| 2020 | Non-linear calibration optimisation based on the Levenberg-Marquardt algorithmabstractAn outstanding calibration algorithm is the most important factor that affects the precision of attitude measurement. This study proposes a non‐linear optimisation algorithm to refine the solutions of the initial guess obtained using the Zhang's technique, the Bouget's technique, or the Hartley's algorithm. Large sets of point correspondences were adopted to test the validity of the proposed method. Extensive practical experiments demonstrated that the proposed method can significantly improve the accuracy of calibration and ultimately obtains higher measurement precision. The error of the reprojection in the proposed method was <0.13 px. At a range of 1 m, the error rate was 0.5% for the length test and about 3% for the angle test. This study proposes a new method to calibrate the relationship between laser radar and the camera. Binocular vision was used to reconstruct the point cloud of the non‐cooperative target. At the same time, data was also obtained using laser radar. Finally, the two groups of systems were fused. Accurate and dense three‐dimensional information of the target was obtained. It could not only obtain the dense pose information of the target surface but also the texture and colour feature information of the target surface. Guoliang Hu, Zuofeng Zhou, Jianzhong Cao |
IET Image Process. | 3 |
| 2020 | Highly accurate 3D reconstruction based on a precise and robust binocular camera calibration methodabstractThe precision of the camera calibration is one of the key factors that affect attitude measurement accuracy in many computer vision tasks. This study proposes a new calibration approach for binocular cameras. Firstly, based on singular value decomposition, the best transformation matrix to the essential matrix is approximated as the initial guess, which is solved in using the Frobenius norm. Secondly, the initial guess is refined through maximum likelihood estimation. A new calculating expression is derived for computing the relative position matrix of the binocular cameras. The Levenberg–Marquardt algorithm is then implemented to refine the initial guess. Large sets of synthesised and real point correspondences were tested to demonstrate the validity of the proposed method. Extensive experiments demonstrated that the proposed method outperforms the state‐of‐the‐art methods. The error rate of the proposed method was 0.5% for the length test and about 1% for the angle test at a range of 1 m. This method can advance three‐dimensional (3D) computer vision one additional step from laboratory environments to real‐world use. Guoliang Hu, Zuofeng Zhou, Jianzhong Cao |
IET Image Process. | 3 |
| 2020 | End-to-end learning interpolation for object tracking in low frame-rate videoabstractIn many scenarios, where videos are transmitted through bandwidth‐limited channels for subsequent semantic analytics, the choice of frame rates has to balance between bandwidth constraints and analytics performance. Faced with this practical challenge, this study focuses on enhancing object tracking at low frame rates and proposes a learning Interpolation for tracking framework. This framework embeds an implicit video frame interpolation sub‐network, which is concatenated and jointly trained with another object tracking sub‐network. Once a low frame‐rate video is an input, it is first mapped into a high frame‐rate latent video, based on which the tracker is learned. Novel strategies and loss functions are derived to ensure the effective end‐to‐end optimisation of the authors’ network. On several challenging benchmarks and settings, their method achieves a highly competitive tradeoff between frame rate and tracking accuracy. As is known, the implications of interpolation on semantic video analytics and tracking remain unexplored, and the authors expect their method to find many applications in mobile embedded vision, Internet of Things and edge computing. Liqiang Liu, Jianzhong Cao |
IET Image Process. | 2 |
| 2020 | Real-time long-term tracker with tracking-verification-detection-refinement
Jiawen Liao, Chun Qi, Jianzhong Cao, Long Ren, Gaopeng Zhang |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Collect and Select: Semantic Alignment Metric Learning for Few-Shot LearningabstractFew-shot learning aims to learn latent patterns from few training examples and has shown promises in practice. However, directly calculating the distances between the query image and support image in existing methods may cause ambiguity because dominant objects can locate anywhere on images. To address this issue, this paper proposes a Semantic Alignment Metric Learning (SAML) method for few-shot learning that aligns the semantically relevant dominant objects through a "collect-and-select'' strategy. Specifically, we first calculate a relation matrix (RM) to "`collect" the distances of each local region pairs of the 3D tensor extracted from a query image and the mean tensor of the support images. Then, the attention technique is adapted to "select" the semantically relevant pairs and put more weights on them. Afterwards, a multi-layer perceptron (MLP) is utilized to map the reweighted RMs to their corresponding similarity scores. Theoretical analysis demonstrates the generalization ability of SAML and gives a theoretical guarantee. Empirical results demonstrate that semantic alignment is achieved. Extensive experiments on benchmark datasets validate the strengths of the proposed approach and demonstrate that SAML significantly outperforms the current state-of-the-art methods. The source code is available at https://github.com/haofusheng/SAML. Fusheng Hao, Fengxiang He, Jun Cheng 0002, Lei Wang 0018, Jianzhong Cao, Dacheng Tao |
ICCV | 5 |
| 2019 | Local difference-based active contour model for medical image segmentation and bias correctionabstractThis study proposes a local bias field and difference estimation (LBDE) model for medical image segmentation and bias field correction. Firstly, the LBDE model uses a linear combination of a given set of smooth orthogonal basis functions, which is called Chebyshev polynomial, to estimate the bias field. Then, a clustering criterion function is defined by considering the difference between the measured image and approximated image in a small region. By applying this difference in the local region, the LBDE model can obtain accurate segmentation results and estimation of the bias field. Finally, the energy functional is incorporated into a level set formulation with a regularisation term, and it is minimised via the level set evolution process. The LBDE model first appears as a two‐phase model and then extends to the multi‐phase one. Extensive experiments on medical images demonstrate that the LBDE model achieves more precise segmentation results in terms of Jaccard similarity and dice similarity coefficient than the comparative models. Therefore the proposed model can increase the segmentation accuracy and robustness to noise. Yuefeng Niu, Jianzhong Cao |
IET Image Process. | 2 |
| 2018 | Non-rigid point set registration by high-dimensional representationabstractNon‐rigid point set registration is a key component in many computer vision and pattern recognition tasks. In this study, the authors propose a robust non‐rigid point set registration method based on high‐dimensional representation. Their central idea is to map the point sets into a high‐dimensional space by integrating the relative structure information into the coordinates of points. On the one hand, the point set registration is formulated as the estimation of a mixture of densities in high‐dimensional space. On the other hand, the relative distances are used to compute the local features which assign the membership probabilities of the mixture model. The proposed model captures discriminative relative information and enables to preserve both global and local structures of the point set during matching. Extensive experiments on both synthesised and real data demonstrate that the proposed method outperforms the state‐of‐the‐art methods under various types of distortions, especially for the deformation and rotation degradations. Zuofeng Zhou, Jianzhong Cao |
IET Image Process. | 3 |
| 2017 | A Novel PDE-Based Single Image Super-Resolution Reconstruction MethodabstractFor applications such as remote sensing imaging and medical imaging, high-resolution (HR) images are urgently required. Image Super-Resolution (SR) reconstruction has great application prospects in optical imaging. In this paper, we propose a novel unified Partial Differential Equation (PDE)-based method to single image SR reconstruction. Firstly, two directional diffusion terms calculated by Anisotropic Nonlinear Structure Tensor (ANLST) are constructed, combing information of all channels to prevent singular results, making full use of its directional diffusion feature. Secondly, by introducing multiple orientations estimation using high order matrix-valued tensor instead of gradient, orientations can be estimated more precisely for junctions or corners. As a unique descriptor of orientations, mixed orientation parameter (MOP) is separated into two orientations by finding roots of a second-order polynomial in the nonlinear part. Then, we synthesize a Gradient Vector Flow (GVF) shock filter to balance edge enhancement and de-noising process. Experimental results confirm the validity of the method and show that the method enhances image edges, restores corners or junctions, and suppresses noise robustness, which is competitive with the existing methods. Jianzhong Cao, Zuofeng Zhou, Jijiang Huang |
Int. J. Pattern Recognit. Artif. Intell. | 2 |